Language is hard. This truth is, in part, why I believe we’ve become so fascinated with generative AI (GenAI) and specifically large language models (LLMs).
Claims range from a greater than 10% chance that AI “could kill all humans” within the next decade to those who consider the technology “AI snake oil.” The latter claim is backed by a world-renowned researcher and empirical data showing models cannot reason or replace your job.
Whatever you believe, the truth likely falls somewhere between these opposing viewpoints.
Caption: Listen to insurance experts discussing their outlooks on the industry (and how you can respond) in the latest episode of Brewing Curiosity – Insurance Unfiltered.
What does our research show?
Unsurprisingly, SAS has found that the largest companies on the planet are using AI technology. In a new SAS report developed with research insights from IDC, 76% of respondents said they “trust” generative AI. Yet 79% of those same respondents will override AI in more than 10% of cases.
Compared to last year’s findings, trust in generative AI has increased significantly. Our 2025 benchmark demonstrated a 200% higher trust in GenAI than in machine learning.
For insurers, we saw great progress across three dimensions:
- The highest increase of any industry in its trustworthiness score.
- The only industry to cross the threshold of perceived to actual trust in AI systems – that is, the Responsible AI capability score has crossed zero, ahead of perceived trust.
- The starkest gap of strong ROI expectations between leaders and laggards.
“Rage” and other takeaways from an NLP expert
The study’s findings echo the truths I heard in a recent conversation with Mary Osborne, SAS’ resident natural language processing expert. Osborne spoke in our Brewing Curiosity: Insurance Unfiltered series, which explores different aspects of today’s data and AI technologies that apply to the insurance industry.
Here, Osborne reacted strongly to the imbalance we see unfolding in real time – too much faith being placed in non-deterministic AI capabilities, namely, large language models.
She uses the term “rage” – an appropriate response considering the ire the technology draws in today’s media coverage. However, the tools hold great potential, and our research shows organizations on the frontier are beginning to see strong ROI returns (about $2 for every $1 invested).
So, what separates the leaders from the laggards? Why do some organizations see results while results elude others? Osborne had good guidance on the how and why.
Read our explainer to learn about natural language processing – see how it works, get examples of technology components, and discover how it’s used today.In her own words
Before we could discuss why some AI investments succeed while others fail, Osborne explained the distinctions between NLP – a branch of artificial intelligence – and generative AI, specifically large language models.
“Natural language processing is the part of AI that helps computers understand, interpret and generate human language – things like text and speech,” she said.
“At a high level, it works by teaching machines to recognize patterns in large amounts of language data. We convert words into numbers, feed millions or billions of examples into models, and those models learn statistical relationships. For example, it learns which words tend to appear together, how sentences are structured and how meaning changes with context.”
On the other hand, Osborne continued, “language models predict the next word based on preceding words. The foundational algorithms for this were developed in the 1940s by a guy named Claude Shannon.”
Claude Shannon is considered a pioneer in information technology. He built one of the earliest examples of machine learning called the “Theseus mouse,” a mechanical mouse that could solve a maze on its own. In 1950, his work opened the door for modern-day AI, even though he did not actually coin the term “artificial intelligence.”
Osborn explained that “LLMs supercharged the ideas Shannon introduced with more data, more parameters and more computational power. ChatGPT, Gemini and Claude (named after Claude Shannon, by the way) are all examples of user-friendly LLMs.”
Want to learn more about GenAI and agentic AI? Check out our AI literacy series, where you can work at your own pace to complete the material. – Learn more about the course.Cause for concern
When considering different branches of artificial intelligence, a common categorization is deterministic versus non-deterministic. The trust being placed in these systems can be somewhat misplaced if the outcomes are not repeatable, as Osborne explored.
“There’s so much hype around LLMs, and most organizations are eager to try them without understanding the downsides. First, they’re non-deterministic –they generate content that isn’t explainable or repeatable. You can ask these models the same question five times, and you aren’t guaranteed to get the same answer.”
In fact, independent research shows that the order of the information in a prompt impacts the output through the primacy and recency effects. If you were to provide an LLM with four options, then ask the same question with the same options but in a different order, there’s a measurable difference in the output.
While this inconsistency is concerning, data privacy issues are troubling as well.
“I’ve talked to people who have inadvertently shared customers’ personally identifiable information (PII) with public models. That’s one of the more insidious things that can happen. Once someone makes a mistake like that (we call it contaminating the model), there’s no easy way to undo it.”
“Treat AI outputs as decision support, not decision makers – especially in high-stakes workflows like underwriting, claims or fraud investigations.”
Mary osborne, SAS
The way forward: Avoiding downside risk
A common question that leaves industry leaders struggling is how to safely use the technology without creating substantial downside risk. Osborne shared her thoughts on that as well.
“A big first step is treating AI outputs as decision support, not decision makers – especially in high-stakes workflows like underwriting, claims or fraud investigations. That means building human-in-the-lead review for important decisions and clearly labeling where AI is providing suggestions versus final answers.
Another important piece is transparency – requiring systems to show why they produced a recommendation, what data sources were used and how confident the model is.”
The human-in-the-lead mindset is evolving among decision-makers and business leaders. In our recent report, SAS found that among insurance leaders – those leading the industry in AI adoption and realizing ROI gains – 78% have cascaded a responsible AI policy to every employee versus just 3% of the laggards.
Osborne’s own commentary underscores this commitment. “Training employees and customers on what these systems can and can’t do goes a long way. The goal isn’t blind automation – it’s augmented intelligence that makes people better at their jobs.”
The final word
We’ll close with Osborne’s final piece of advice.
“If you’re implementing LLMs, don’t throw out the best practices you’ve always used for other types of models. Make sure you have a solid strategy for evaluating the output so you have a clear understanding of the model's actual performance. If you understand how well the models are performing, you can adjust to continually improve the user experience.”
Check out this episode in our series, and stay curious, my friends!