There is a particular kind of business problem that does not announce itself. It accumulates. A system that was built to handle one thing is gradually asked to handle adjacent things, then edge cases, then exceptions. Nobody signs off on the expansion. It simply happens — through a helpful colleague adding a new input, a manager broadening the prompt, or an enthusiastic team member connecting a new data source. Six months later, the system is doing four times the work it was designed for, and nobody has checked whether it is still doing any of it well.

This is scope drift, and it is currently the most underappreciated failure mode in enterprise AI. It is not dramatic. There is no crash, no error message, no obvious moment of failure. The system keeps producing outputs. It just produces them with less accuracy, less relevance, and less reliability than it did on day one — and because the change is gradual, it is easy to miss.

The question worth sitting with is not whether your AI system was built well. It probably was, for the task it was originally given. The question is whether the task it is actually doing today bears any resemblance to the one it was designed for.

Why scope drift happens

AI systems — particularly those built on large language models — are generalist by nature. They do not refuse work they were not designed for. They attempt it. This is what makes them useful and what makes them dangerous when left unsupervised. A system built to summarise incoming client communications will, if asked, also attempt to classify them, prioritise them, draft responses to them, and flag sentiment trends. It will do all of this without complaint and without any indication of where its reliability ends.

The business logic that drives expansion is usually sound. If the system handles task A well, why not give it task B? The problem is that the system's actual competence — its accuracy, its calibration, its appropriate use of the information it has access to — was validated only against task A. Task B may look similar from the outside and be fundamentally different in the ways that matter. A system trained and tested on structured client queries may perform poorly on unstructured internal communications. A system that summarises accurately may classify poorly. The boundaries are rarely obvious.

There is a further complication. As an AI system's working context fills with more varied inputs, its ability to maintain coherence across all of them degrades. Research published in 2025 and 2026 consistently finds that context quality — the ratio of relevant, current information to noise — matters more than context size. A system asked to handle ten loosely related tasks simultaneously is not ten times as capable as a system handling one task well. It is likely less capable at all of them.

The signal you are probably missing

Scope drift rarely produces visible errors. It produces invisible ones. The system continues to function. Outputs continue to arrive. The problem is that the outputs are being evaluated against the wrong standard — what the system used to do, rather than what it should be doing now.

There are, however, early signals worth watching for. The first is output homogeneity: the system begins producing responses that are structurally similar regardless of what was asked. This is a sign that it is defaulting to its most familiar pattern rather than reasoning carefully about each input. The second is escalation silence: the system stops flagging uncertainty or routing edge cases for human review, not because it has become more capable, but because its calibration has drifted and it no longer recognises what it does not know. The third is input creep: the volume, variety, or format of inputs has changed materially since the system was validated, but no one has re-tested it against the new input profile.

The uncomfortable truth is that most organisations do not have a formal process for detecting any of these signals. They have a deployment process, and they have a feedback mechanism for obvious failures. They do not have a systematic way of asking whether the system's remit has quietly expanded beyond its designed capability.

What good governance actually looks like

The answer is not to lock systems down so tightly that they cannot be useful. Useful AI systems need room to operate. The answer is to be deliberate about how scope is extended, and to build the monitoring infrastructure that makes expansion visible.

Practically, this means three things. First, a maintained record of what the system was originally validated to do — its intended inputs, its intended outputs, and the conditions under which it was tested. This is not complex documentation; it is a one-page description of the system's remit. Most organisations do not have it. Second, a periodic review — quarterly is usually sufficient — that compares current usage against original intent and asks explicitly whether the gap has grown. Third, a defined escalation path: a clear answer to the question of what the system should do when it encounters something outside its validated scope, whether that is flagging for human review, declining to proceed, or routing to a different process.

This last point deserves emphasis. A well-designed AI system is not one that attempts everything it is given. It is one that knows the boundary of its reliable operation and behaves differently at that boundary. The 2026 Singapore Consensus on Global AI Safety Research Priorities frames this precisely: human oversight should involve structured decisions at defined intervention points, calibrated to the risk and reversibility of the action, rather than continuous supervision of every step. That framing applies equally to commercial AI deployments. You are not trying to watch everything. You are trying to ensure that consequential uncertainty reaches a human being.

The remit review: three questions worth asking this week

You do not need a formal audit to begin. Three practical questions will surface most of the risk.

  • What is the system doing today, in concrete terms? List the actual inputs it receives, the outputs it produces, and the decisions or actions those outputs feed into. Compare that list with the original design intent. If you cannot produce the original design intent, that absence is itself the answer.
  • Has the input profile changed since the system was last validated? Input changes are the most common and least monitored source of scope drift. New data sources, new team members, new business processes, different languages or formats — any of these can shift the system's effective operating conditions without anyone noticing.
  • What happens when the system encounters something it should not handle? If the answer is that it handles it anyway, you have an escalation design problem. The system should have a defined behaviour at the edge of its competence: flag, pause, route, or decline. If it has none, it is operating without guardrails in precisely the situations where guardrails matter most.

The deeper issue: trust that was never re-earned

The reason scope drift is so commercially significant is that it erodes trust without anyone noticing it is being eroded. Teams continue to rely on the system's outputs. Decisions continue to be made on the basis of those outputs. But the underlying reliability has quietly degraded, and nobody has gone back to check whether the trust the organisation places in the system is still warranted.

McKinsey's 2025 State of AI survey found that while 88 percent of organisations now use AI in at least one business function, only around a third have begun to scale their programmes in any systematic way. The gap between adoption and mature deployment is wide. Part of what fills that gap is exactly this: organisations that have deployed AI systems without building the operational discipline to maintain them as the work evolves around them.

The organisations that use AI well do not necessarily have more sophisticated systems. They have cleaner boundaries. They know what each system is for, they can see when that is changing, and they have a process for deciding whether the change is acceptable. That is not a technology problem. It is a management problem, and it is more tractable than most.

A closing thought

The most durable AI systems in commercial use are not the ones with the broadest capabilities. They are the ones with the clearest remit, the most honest accounting of where their reliability ends, and the most deliberate process for extending their scope when the business requires it.

If there is a recurring process in your organisation where AI outputs feed consequential decisions — and most organisations now have several — it is worth asking when someone last checked whether the system is still doing the job it was designed for. Not whether it is producing outputs. Whether the outputs remain fit for purpose.

If that question surfaces something worth thinking through, we are happy to discuss it.

Before we talk.

You do not need a solution in mind. Bring one recurring bottleneck, missed signal or decision that should work better.

Start a conversation