There is a question that tends to surface three to six months after an AI system goes live in a business. It rarely comes from the technology team. It comes from a founder, a department head, or a client who notices something unexpected: the system made a decision — or sent a communication, or filtered a list, or declined a request — that nobody intended it to make.

The question is not 'why did it do that?' The real question, once you dig into it, is: 'Who told it what it was allowed to do?'

In most cases, the honest answer is that nobody told it clearly. A prompt was written. A workflow was assembled. Some parameters were set. And then the system was handed a set of real tasks with real consequences. The boundaries were assumed rather than designed.

This is not a rare edge case. It is the standard condition of AI deployment in commercially active organisations right now. And it is worth addressing not because it creates dramatic failures — though it sometimes does — but because it quietly degrades the quality of every output the system produces.

The mandate problem

Every AI system that operates inside a business has an effective mandate: the set of things it will do, the way it will prioritise between competing options, the situations in which it will act without prompting, and the thresholds at which it will pause or defer. Whether or not anyone wrote this mandate down explicitly, it exists. The system behaves as if it does.

The problem is that when a mandate is not written down, it is filled in by defaults — the model's own training, the way the prompt happens to be phrased, the first few edge cases that got resolved informally, and whatever assumptions were baked in during setup. These defaults are not necessarily wrong. But they are not yours. They were not chosen deliberately for your context, your clients, or your risk tolerance.

The distinction matters enormously. A mandate that was designed tells the system what to optimise for, what to protect, what to escalate, and what to refuse. A mandate that was assumed tells the system nothing of the sort — it simply inherits whatever the path of least resistance happened to produce.

Most organisations cannot point to a document that answers four basic questions about each AI system they operate: What is it for? What is it not for? What may it do without human review? And who is accountable when it behaves unexpectedly? If those questions do not have clear answers, the mandate is assumed, not designed.

Why this is becoming a structural issue

For a long time, the assumed-mandate problem was manageable because AI systems had relatively narrow roles. A system that drafted email subjects or scored leads could drift slightly from its intended behaviour without much consequence. The outputs were easy to inspect. The stakes were limited.

That situation has changed. AI systems are increasingly asked to take actions — sending communications, routing enquiries, filtering candidates, adjusting pricing, managing workflows — rather than merely producing outputs for human review. The gap between what the system does and what a person verifies has widened considerably.

This shift has a regulatory dimension too. The EU AI Act's high-risk AI obligations, which came into full enforcement in August 2026, require deployers of qualifying systems to assign human oversight formally, retain logs, and operate systems in accordance with documented instructions. The regulation is specific about accountability: if your organisation uses a high-risk AI system, you carry statutory obligations regardless of whether you built the system yourself or purchased it as a service.

Even for organisations whose systems do not meet the regulatory threshold for 'high-risk', the underlying principle is sound. If you cannot describe what your AI system is authorised to do, you are not in a position to say whether it is doing it.

What a designed mandate looks like

A mandate is not a long document. It is a set of deliberate decisions, ideally recorded in a form that can be reviewed and updated. The decisions fall into four areas.

The first is scope. What tasks is the system designed to handle, and — equally important — what tasks are explicitly outside its remit? A system with no explicit exclusions will attempt to be helpful wherever it can, which sounds positive until it produces an output in an area where it has no reliable knowledge or no authority to act.

The second is authority. Which outputs may the system deliver without human review? Which require a person to confirm before anything happens? Where is the line drawn, and is that line appropriate given the consequences of an error? Systems that act on behalf of a business — communicating with clients, updating records, triggering workflows — need their authority level stated explicitly, not inherited from the path of least resistance.

The third is escalation. Under what conditions should the system pause and refer the matter to a person? This is not only a safety question. A well-designed escalation condition is also a quality filter: it identifies the situations in which the system is operating outside its reliable range and routes those situations to someone equipped to handle them.

The fourth is accountability. When a system produces an unexpected outcome, who is responsible for investigating, deciding how to respond, and updating the system's behaviour? If that accountability is diffuse — spread across whoever built it, whoever uses it, and whoever commissioned it — then nothing will be corrected systematically.

The context problem underneath the mandate

There is a technical layer to this that is worth understanding without needing to go deep into the engineering. Every AI system reasons over a defined window of information at any given moment — the data and instructions it has been given for that specific task. What sits outside that window is, in a practical sense, invisible to the system.

This creates a predictable failure mode: a system may behave impeccably in the situations it was tested on, then produce a poor output in a situation that is superficially similar but differs in a way that matters. Not because the system is malfunctioning, but because the information that would distinguish the two situations was not in the window it was working from.

A mandate that was designed accounts for this. It does not assume the system will always know what it needs to know. It builds in checkpoints: situations in which the system is instructed to check whether it has sufficient context before proceeding, rather than producing a confident output based on an incomplete picture.

When a system operates without those checkpoints — producing polished, fluent outputs regardless of how much it actually knows about a given situation — the problem is not visible in any single output. It accumulates silently across many interactions. The signal that something is wrong is often diffuse: outputs that are technically correct but slightly off, recommendations that miss something contextual, communications that do not quite fit the situation.

How to approach this practically

The starting point is not a technology audit. It is a conversation about accountability. For each AI system operating in your business — whether it was built bespoke, assembled from tools, or purchased as part of a platform — the useful questions are straightforward.

  • What was this system originally designed to do, and has that scope expanded since deployment?
  • What can it do without a person reviewing the output first?
  • When did someone last check whether its behaviour still matches the intent behind it?
  • If it produces an unexpected or incorrect output today, who is responsible for addressing that?
  • Is there a record of what it is authorised to do, or is that knowledge held informally by whoever built it?

The review cadence

A mandate that is defined once and never revisited is only marginally better than no mandate at all. The environment a system operates in changes: the data it receives shifts, the tasks it is asked to handle expand, the regulatory context evolves, and the organisation's own risk appetite may not be the same as it was six months ago.

Sensible practice is to build a review cadence that is proportional to the consequences of the system's outputs. A system that drafts internal summaries might need reviewing quarterly. A system that communicates directly with clients or influences commercial decisions warrants closer and more frequent attention.

The review does not need to be technically deep. The most important questions are organisational, not algorithmic: Is this still doing what we intended? Has its scope drifted? Do the people using it understand its limits? And does the accountability for its outputs sit with someone who has both the authority and the information to act on a problem?

A note on proportionality

None of this requires a formal governance programme. For most SMEs and leadership teams, the investment involved in defining and reviewing an AI mandate is small relative to the cost of discovering — usually at an inconvenient moment — that a system has been operating outside its intended parameters for some time.

The organisations that handle this well are not necessarily the ones with the most sophisticated AI. They are the ones that treat the mandate as a management decision rather than a technical setting — something that belongs in the same category as deciding who can approve a contract or sign off on a proposal, not something that was answered implicitly when the system was set up.

If there is a recurring bottleneck in how your AI systems are defined, reviewed, or held to account, that is usually worth a direct conversation. The answers are rarely complicated — but they do need to be answers, not assumptions.

Before we talk.

You do not need a solution in mind. Bring one recurring bottleneck, missed signal or decision that should work better.

Start a conversation