There is a moment in almost every automated process where something unexpected appears. A contract term that does not fit the template. A client response that sits outside the range the system was trained to handle. A number that is technically within tolerance but feels wrong to anyone who knows the account. In well-designed systems, that moment triggers a clean handover. In most systems deployed today, it triggers nothing — and the automation continues.
That continuation is where the damage accumulates. Not in dramatic failures that everyone notices, but in the quiet compounding of outputs that were slightly off, decisions that should have had a second pair of eyes, and actions taken at speed because nobody had drawn a line and said: this far, then stop.
The question of when an AI system should yield — pause, flag, or hand the task to a human — is not a technical question. It is a design question. And most organisations building or buying AI capability are not answering it deliberately.
Automation does not fail where you expect it to
The common assumption is that AI systems fail on inputs they cannot understand — garbled data, missing fields, requests entirely outside their scope. Those failures are visible. The system produces an error, a blank output, or something obviously broken. Someone notices.
The failures that cost more are the ones that look fine. A system that handles ninety percent of cases well will, given sufficient volume, produce consequential errors in the remaining ten percent — and produce them with the same surface confidence it brings to everything else. There is no flashing light. The output arrives, is processed by the next step in the workflow, and the error becomes embedded.
This is the structural problem that escalation design exists to solve. The goal is not to reduce automation — it is to define, before anything goes live, the conditions under which the system should stop trusting itself.
Two different kinds of stop
It helps to distinguish between two situations that are often treated as the same thing.
The first is a confidence stop. The system encounters an input it cannot assess reliably — the data is ambiguous, contradictory, or simply outside the distribution it was designed for. A well-constructed system will signal this explicitly rather than guess. Designing for this means defining thresholds: below a certain confidence level, the task routes to a human rather than completing automatically.
The second is a consequence stop. The system is confident enough in its assessment, but the action it is about to take is sufficiently consequential — irreversible, high-value, or with significant downstream effects — that confidence alone should not be sufficient authorisation. A contract clause accepted. An invoice submitted. A client communication sent. A procurement decision made. These are actions where the potential cost of error is high enough that a human checkpoint is warranted regardless of how certain the system appears to be.
Most escalation thinking stops at the first category. Organisations ask: what does the system do when it is unsure? They rarely ask: what does the system do when it is sure but wrong about something that matters? That second question is the harder one, and it is the one that shapes whether an AI deployment is genuinely trustworthy in practice.
The cost of getting the handover wrong in both directions
There is a temptation to solve this by escalating everything significant. If in doubt, route it to a human. The instinct is cautious and understandable, but it creates its own problem.
When reviewers are asked to approve too much, they stop reviewing. The workload of checking high volumes of AI outputs — many of which are entirely routine — produces what practitioners sometimes call automation bias: the tendency to ratify what the system has produced without genuinely interrogating it. The oversight exists on paper. In practice, consequential errors pass through because the reviewer's attention has been exhausted on the trivial ones.
The opposite failure is equally costly. A system given too much autonomy, without clear thresholds for when it should pause, will eventually take an action in a live environment that produces damage at a speed no manual process could match. Automated systems operate faster than human review cycles. When they go wrong without a stop condition, the error is replicated across every subsequent step before anyone realises what has happened.
Getting this right is not primarily a technology problem. It is a governance problem: who decides where the lines are drawn, on what basis, and who reviews those decisions as the system's environment changes over time.
What a well-designed escalation point looks like
A useful mental model is to think about escalation as a structural feature of the system, not an emergency override. The distinction matters. Emergency overrides are things people activate when something has already gone wrong. Structural escalation points are built into the normal operating path — they activate before the error occurs, not after it is discovered.
Practically, this means making three decisions before any system goes into operation. First: what are the categories of action this system should never take without human sign-off, regardless of confidence? These are defined by consequence, not complexity. Second: below what confidence threshold should the system stop and route rather than complete? This requires an honest assessment of the cost of being wrong versus the cost of the delay a human review introduces. Third: who receives the escalation, with what context, within what time window, and what happens if they do not respond?
That last point is more often neglected than the first two. An escalation path that routes to a shared inbox, with no defined owner, no time limit, and no stated consequence for non-response, is not an escalation path. It is a way of making the organisation feel it has addressed the problem without actually doing so.
- Define escalation by consequence, not just by complexity or confidence.
- Distinguish between actions that require human approval and actions that require human awareness.
- Specify who receives the escalation, not just which team or function.
- Set a response window and a default outcome if the window lapses.
- Review the thresholds at regular intervals, particularly after any significant change in volume, context or system capability.
The handover itself is a design problem
Even when organisations define their escalation thresholds correctly, they frequently underinvest in the quality of the handover. The human reviewer receives a flag that something requires attention, but what they receive alongside it — the context that would allow them to make a good decision quickly — is thin or absent.
A reviewer who must reconstruct the background of a case before they can assess what the system has done is not a meaningful check. They are a bottleneck. The overhead of the human review becomes large enough that pressure builds to reduce the number of escalations, which typically means raising the thresholds — which means more passing through without review. The design problem compounds.
What the reviewer needs at the point of escalation is: what the system was trying to do, what it found, what action it was about to take, why it stopped, and what the time sensitivity of a response is. If that information cannot be surfaced cleanly at the moment of handover, the escalation architecture has not been completed. It has only been started.
The regulatory dimension is becoming harder to ignore
For organisations operating in regulated sectors — financial services, healthcare, legal, anything touching personal data at scale — the question of escalation design is moving from best practice to compliance requirement. The EU AI Act requires demonstrable human oversight in high-risk applications, not nominal oversight. The difference is significant: demonstrable means documented thresholds, logged escalations, evidence that reviewers had sufficient information and authority to act, and records of what decisions were made and by whom.
Even for organisations not directly subject to these frameworks, the precedent they set is shaping expectations more broadly. Clients, counterparties and auditors are beginning to ask questions about how AI systems are governed that would not have been asked two years ago. Having a coherent answer — one that goes beyond 'there is a human in the loop' — is increasingly a commercial consideration, not only a risk management one.
A question worth sitting with
For any AI system currently operating within your business, or any you are considering deploying, the most useful question to ask is not 'what can this system do?' It is: 'Under what conditions has this system been designed to stop?'
If the answer is clear — if someone can describe the specific thresholds, the consequence categories, the escalation recipients, the handover context and the review cadence — then the system is governed. If the answer is vague, or if the question produces a pause, then what has been deployed is a system whose boundary conditions have not been deliberately designed. That is a different kind of risk than most technology conversations surface.
The stopping point is not the failure mode. It is the feature that makes everything else trustworthy. If you have a recurring workflow where the boundaries of AI action have never been formally drawn, that is usually a productive place to start a conversation.
Before we talk.
You do not need a solution in mind. Bring one recurring bottleneck, missed signal or decision that should work better.
Start a conversation