Client case study · AI in Business · LLM Production

The
Production
Paradox.

88% of organisations now use AI. 74% have rolled back at least one AI agent after it went live. I analysed three independent industry reports to find out why the gap between deployment and survival is so wide — and what actually separates the systems that stay.

88%Use AI in at least one function (McKinsey)
74%Rolled back ≥1 agent after production (Sinch)
3.3%Lowest hallucination rate on the leaderboard (Vectara)
6×Gap between best and worst model (Vectara)

01 —
The Adoption
Gap.

McKinsey's State of AI 2025 (1,993 respondents, 105 countries) captures a market at full tilt. 88% of respondents report using AI in at least one business function. 62% are experimenting with or already running AI agents. 23% are scaling at least one agentic system. And 39% report some measurable impact on EBIT.

Use AI in ≥1 function
88%
Experimenting / using agents
62%
Scaling ≥1 agentic system
23%
Report AI impact on EBIT
39%
The gap is between "using" and "scaling." 88% use AI but only 23% are scaling an agentic system — a 65-point drop. Adoption is near-universal; durable, scaled deployment is rare. The reports show the reason is not the model. It's what happens between the demo and production.

02 —
The Production
Paradox.

Sinch's AI Production Paradox (2,527 business leaders, 10 countries) measures what happens after deployment. 62% of organisations have an AI agent running in production on customer-facing channels. But 74% of those organisations rolled back or restricted at least one agent after it went live.

Agent in production
62%
Rolled back ≥1 agent
74%

Why agents get pulled

Data / privacy leak
31%
Hallucination / brand risk
22%
Couldn't diagnose the fault
16%
81% of rollbacks happen at organisations with mature control mechanisms. This is the paradox: the teams that test most rigorously roll back most, because their controls surface more problems. Good governance doesn't reduce rollbacks — it makes them visible. The other finding: 84% of engineering teams spend at least half their time building safety infrastructure ("guardrail tax") instead of product features. The cost of failure lands in three places: support queue (35%), brand (34%), and engineering time (the 84% guardrail tax).

03 —
The Hallucination
Benchmark.

Vectara's next-generation Hallucination Leaderboard (7,700 articles, up from 1,000) measures how often models introduce hallucinations when summarising. There is no single "safe" rate for business chatbots. The spread runs from 3.3% (Gemini 2.5 Flash-Lite) to over 10% for reasoning models — a 6× gap. And model size is not the deciding factor: the leaderboard's top tier mixes 32B+ and sub-32B models.

Finix S1 32B (best)
1.8%
GPT-5.4 Nano
3.1%
Gemini 2.5 Flash-Lite
3.3%
Phi-4
3.7%
Llama 3.3 70B
4.1%
Mistral Large 2411
4.5%
Nova Pro v1
5.1%
Gemma 3 27B
7.4%
GPT-5.5
9.3%
Claude Sonnet 4
10.3%
GPT-5.2 High
10.8%
DeepSeek-R1
11.3%
Reasoning models hallucinate more, not less. The "smart" reasoning models (GPT-5.5, Claude Sonnet 4, DeepSeek-R1) sit at 9–11%, while smaller non-reasoning models like Gemini 2.5 Flash-Lite (3.3%) and Phi-4 (3.7%) are far safer. For a business chatbot, the benchmark pick is a small grounded model, not the flagship.

Four root causes of business hallucination

Vectara · Leaderboard

RAG retrieval failure

Hallucination rate scales with article complexity and document length — long, technical documents produce more hallucinations than short news items. The problem is the retrieval layer, not the model.

Sinch · Production Paradox

No production-grade testing

74% rollback rate is direct evidence. 81% of rollbacks hit organisations with mature controls — meaning problems exist but were caught late. Models ship without validation in near-production conditions.

Vectara · Leaderboard

Wrong model for the task

Reasoning models score >10% hallucination; the best small models hit 3.3%. A 6× gap means model choice is decisive — but size alone doesn't predict quality.

Sinch · Engineering Survey

Guardrail tax

84% of engineering teams spend ≥50% of their time on safety infrastructure instead of product. Teams fight the symptom (hallucination, rollback) rather than the root cause, because the right tooling isn't available.

04 —
Chatbot Classes
&Cost.

Based on the three reports, business chatbots fall into four classes. The higher the class, the more capability — and the more engineering time, maintenance, and rollback risk. Cost is measured in developer hours (deployment) and the administrative-time equivalent of ongoing maintenance (a full-time EU admin role = 160 h/month).

Chatbot class Deployment
dev time
Maintenance
per month
Call-center FTEs
replaced
1Rule-basedScripted FAQ 1–3 days32–96 dev-hours 4–8 hadmin-time equiv. 1–2
2RAGRetrieval-grounded 2–7 days64–160 dev-hours 16–24 hadmin-time equiv. 3–5
3AgenticTool-using decisions 7–14 days160–280 dev-hours ~40 h≈ 1 admin week 5–10
4Multi-agentOrchestrated agents >14 days>280 dev-hours ~80 h≈ 2 admin weeks 10–20+
Cost climbs linearly. Risk compounds.
  • Class 2 (RAG): 64–160 dev hours, replaces 3–5 call-center seats — the sweet spot where ROI is easiest to prove.
  • Class 4 (multi-agent): 280+ dev hours, replaces 10–20+ seats — but this is exactly the class where the 74% rollback rate and the 84% guardrail tax bite hardest.
  • The ROI only holds if the deployment is validated on your real documents, in your domain, before it reaches production.

What I'd
do
Differently.

  1. Set the expectation, not the model.

    Don't treat the 74% rollback rate as an anomaly — it's the base rate for customer-facing AI. Plan the deployment as if a rollback is likely, and design the recovery path (audit trail, fallback to human) before the first user touches the agent.

  2. Pick the model by domain benchmark, not by name.

    For low hallucination, the leaderboard points to small grounded models (Gemini 2.5 Flash-Lite 3.3%, Phi-4 3.7%), not flagships. Validate the specific model in the target vertical (legal, medical, finance) — general benchmarks understate domain-specific hallucination.

  3. Invest in the audit trail first.

    16% of rollbacks are caused by the inability to diagnose the fault. A full decision log — what the model retrieved, what it generated, which guardrail fired — pays back multiplicatively. It turns a 16% "unknowable" rollback into a fixable one.

  4. Test in the domain, not just the benchmark.

    Hallucination scales with document complexity and length. Run the retrieval pipeline on your actual documents — long, technical, domain-specific — before go-live. A 3.3% rate on news articles can become 10%+ on a legal contract.

  5. Roll out in risk order.

    Start with low-risk, high-volume queries (FAQ, routing) and graduate to high-risk, high-stakes ones (billing, legal advice) only after the low-risk layer has proven stable in production. This keeps the 74% rollback rate from hitting the part of the business that can't afford it.

Sources
&Method.

McKinsey
The State of AI in 2025 · n=1,993 · 105 countries
Adoption data: 88% AI usage, 62% agent experimentation, 23% scaling, 39% EBIT impact, 51% negative effects. Survey conducted June–July 2025.
Sinch
The AI Production Paradox · n=2,527 · 10 countries · 6 industries
Production reality: 62% agents live, 74% rollback rate, 81% of rollbacks at mature-governance orgs, 84% guardrail tax, failure cost split (35% support / 34% brand). Survey January 2026.
Vectara
Next-Gen Hallucination Leaderboard · 7,700 articles · HHEM
Model benchmark: grounded hallucination rates from 1.8% (Finix S1) to 11.3% (DeepSeek-R1). Rate scales with document complexity and length. Updated May 2026.
Method. All three sources are independent industry reports. The hallucination figures come from Vectara's public leaderboard (HHEM evaluation model). The deployment and maintenance cost estimates are my own synthesis, anchored to the Sinch engineering data (guardrail tax, rollback causes) and standard EU administrative-hour references. No single source covers the full picture — the case study is the cross-reference.

Need this
for your
deployment?

I build the evaluation framework, test the models against your real documents, and design the deployment path that survives contact with production. Write to me and we'll turn your AI deployment from a rollback waiting to happen into a system that stays.