Show notes
This episode is based on the Coinmedia article “Why Your AI Agent Works in the Demo and Fails in Production”, linked above. It draws on research published by Gartner in May 2026, and on data from Forrester, Anaconda and BCG published in 2026.
The Gartner prediction referenced in the episode: by 2027, 40 per cent of enterprises will decommission AI agents because their teams never distinguished between an agent's capacity to act and the scope of access it should have been granted. The 88 per cent pilot-to-production failure figure comes from Forrester and Anaconda's 2026 joint research.
The three-tier authority model discussed in the episode is a practical framework for leadership teams, not a technical implementation guide. The point is to prompt the governance conversation before a deployment is committed to, rather than after it has stalled.
Chapters
- 0:08 Cold open
- 0:55 Why pilots always look better than production
- 2:31 The structural mistake underneath the failure rate
- 3:30 A mental model: three tiers of authority
- 5:16 What the organisations that get it right actually do
- 6:52 Close
Transcript
Cold open
Picture a meeting from about a year ago. Someone runs a demonstration of a new AI system. It handles the questions cleanly. It produces exactly the kind of output the team has been asking for. The people in the room are persuaded. The project gets approved.
Then it goes live. And within a few months, something has quietly gone wrong. The system is producing outputs nobody fully trusts. People have started working around it. The project is technically still running, but nobody is really using it. And the people who backed it are beginning to wonder whether the technology was ever as good as it looked.
Here is the thing. In most of those cases, the technology was fine. What wasn't ready was the governance around it. That is what this episode is about — why that gap appears, and what it actually takes to close it.
Why pilots always look better than production
Start with the pilot itself, because that is where the problem is usually seeded.
A pilot is a curated environment. The data going into it has been tidied up. The workflow has been defined in advance. The edge cases — the messy requests, the late inputs, the things that don't quite fit any category — those are either excluded or handled manually by whoever is running the demonstration. None of that is dishonest. It is simply how pilots work. But it means that a successful pilot is measuring performance under ideal conditions, not resilience under normal ones.
Production is the opposite of ideal conditions. Data arrives late, or in the wrong format, or not at all. Users develop habits the designers never anticipated. A request lands that sits just outside the system's defined scope, and somebody has to decide what to do with it. Over time, small deviations accumulate. The gap between what the system was built to do and what the organisation actually needs it to do keeps widening.
Forrester and Anaconda published data this year — twenty twenty-six — showing that eighty-eight per cent of AI agent pilots fail to graduate to production. The blockers people cited were not model quality. They were evaluation gaps, governance friction, and reliability concerns in live conditions. These are solvable problems. But they require a different kind of preparation than most organisations apply to their pilots.
The structural mistake underneath the failure rate
In May of this year, Gartner published research that named the structural mistake driving this pattern. And it is worth sitting with, because it is not obvious until someone points it out.
Most enterprises apply the same governance policy to every AI system they run — regardless of how much those systems actually do. A system that surfaces a ranked list of prospects for a sales team to look at gets the same treatment as a system that sends emails to clients or updates records in a shared database. Gartner's prediction is that by twenty twenty-seven, forty per cent of companies will have decommissioned AI agents specifically because they never made that distinction.
Gartner's analyst framed it plainly: enterprises are treating agent governance as binary. Either the system is locked down, or it is fully trusted. And that binary is the root cause of failure. Because locked down too tightly, the system becomes useless. Trusted too broadly, it becomes dangerous. Neither outcome is what anyone intended.
A mental model: three tiers of authority
So here is a mental model that helps. Think about every AI system in your organisation in terms of what actually happens when it acts. That gives you three distinct tiers, and the governance you need is different for each one.
The first tier is advisory and internally contained. A system analyses something and surfaces a result for a human to review. If the system is wrong, a person catches it before anything external happens. The governance here can be relatively light — periodic review of output quality, a clear owner, a way to flag systematic errors.
The second tier involves actions with external or financial consequences, but within tightly defined parameters. Sending a follow-up to a prospective client. Triggering a payment below a defined threshold. Updating a record that other people depend on. This tier needs explicit authorisation boundaries. Not just a policy document — a technical limit on what the system can and cannot touch. And it needs a named human who is accountable for what the system does in their name.
The third tier is where the system exercises genuine judgement in genuinely ambiguous situations. Negotiating terms. Making exceptions to standard policy. Coordinating across multiple other systems without a human in the loop at each step. This tier exists and it is commercially valuable. But it requires the most careful design, the most explicit guardrails, and a very clear escalation path for when the system reaches the edge of its competence.
Most failures happen because a second or third tier system was governed as if it were first tier. The organisation saw capability in the pilot and assumed that implied trustworthiness in production. Those are not the same thing.
What the organisations that get it right actually do
The picture from this year's research is not uniformly discouraging. The same data that surfaces the forty per cent failure prediction also identifies what separates the deployments that work.
Ownership is the single most consistent differentiator. Among enterprises that have successfully moved AI agents into production, fifty-six per cent now name a dedicated person — an AI agent owner, an agentic ops lead, whatever the title — who is accountable for that system's behaviour the way a product owner is accountable for a product. In twenty twenty-four, that figure was eleven per cent. The organisations that got there were not the ones with the largest budgets. They were the ones that treated the deployment as an operational commitment, not a technology project.
Clarity of scope comes second. B C G and Forrester data puts the median payback period at five point one months across functions, and the fastest paybacks are concentrated in workflows that share three characteristics: well-documented, high-volume, and measurable. Customer service triage, sales development outreach, invoice matching, internal reporting. In each case, the agent is not exercising broad discretion. It is executing a defined process reliably and flagging exceptions cleanly. That is a solvable design problem.
The deployments that struggle have the opposite characteristics. Fuzzy scope, no measurable success criteria, and an assumption that the system will self-organise around whatever the organisation throws at it. Those projects survive the pilot and collapse around six months into production. Exactly the pattern Gartner describes.
Close
If there is one thing to take away from all of this, it is a reframe of the question. Most organisations approach an AI deployment by asking: what can this system do? That is the right question for a procurement conversation. The question that determines whether a system actually survives in production is different. It is: what should this system be permitted to do, in which contexts, with what level of human review before it acts?
Those are governance questions. They require the same rigour a finance director applies to signing authority. Who can commit what, up to what amount, under what circumstances. And they need to be answered before the system goes live — not after the first production incident.
The question to take back to your own team this week is a simple one. For the AI system you are running, or planning to run — which tier of authority does it actually need to be useful? And is the governance you have planned calibrated to that tier, or is it the same policy you apply to everything else?
If that question surfaces something worth talking through — a deployment that has stalled, a pilot that never made it to production, a governance question without a clear owner — you are welcome to get in touch. The details are in the show notes.
Each episode is written from that week's Coinmedia insight and voiced with a synthetic model of Zsófia's own voice. The thinking, the editorial line and the approval to publish are human.
Before we talk.
You do not need a solution in mind. Bring one recurring bottleneck, missed signal or decision that should work better.
Start a conversation