There is a question that almost never appears in a vendor demonstration, a procurement checklist or a board update on AI progress: what does this system actually believe right now, and is that belief current?

The question matters because most operational AI failures are not caused by the model choosing the wrong word or producing a nonsensical output. They are caused by a system acting confidently on information that is incomplete, contradictory or simply no longer true. A pricing agent that does not know a discount policy was revised last quarter. A customer-routing system that does not know a product line was discontinued. An analysis tool that does not know the reporting structure changed after a recent reorganisation.

These are not edge cases. They are the ordinary consequence of deploying a capable system without governing what that system is allowed to treat as fact.

The confidence problem

Language models and AI agents share a characteristic that is genuinely unusual compared with most software: they do not announce uncertainty in the way that a spreadsheet returns an error or a database query returns a null. They produce a response. The response is coherent. It sounds authoritative. It may even be largely correct — which is precisely what makes the exceptions difficult to catch before they cause a problem.

This characteristic becomes commercially significant the moment an AI system is doing something consequential: drafting terms, qualifying a prospect, summarising a contract, recommending a supplier, flagging a risk. In each of those cases the output is only as reliable as the information the system was given to work with.

What a system 'knows' is not fixed at the moment you install it. It changes — or fails to change — as your business changes around it. Policies are revised. People leave. Products evolve. Markets shift. If the context feeding the system does not keep pace, the system's confident outputs increasingly diverge from operational reality.

Context is the hidden variable

When organisations audit an underperforming AI deployment, they tend to examine the model — whether it is powerful enough, whether it was configured correctly, whether a different vendor would produce better results. The model is rarely the problem. The context is.

Context, in practical terms, is everything the system draws on before it responds: the documents it was trained or fine-tuned against, the data it retrieves in real time, the policies and definitions it has been given, the conversation history it carries forward, and the implicit assumptions baked in during setup. Any one of those layers can go stale.

A 2026 benchmark study on AI agent memory found that agents can reuse invalid or obsolete information when the facts they were trained on have since changed — and that in an enterprise setting this might mean acting on a retired policy, a deprecated metric or data from a system that no longer feeds current reporting. The system does not know it is working from an old map. It simply navigates with confidence.

A separate analysis published by Forbes in March 2026 described this plainly: when AI agents act on incomplete business context, the consequences extend beyond a poor answer. They include operational disruption, financial exposure and compliance risk. One illustrative pattern involved a quoting agent generating contracts below approved pricing thresholds — not because the model was broken, but because it had not been told the thresholds had changed.

Three layers leaders rarely inspect

Most leadership conversations about AI governance focus on access controls, data privacy and model behaviour. Those are legitimate concerns. But there are three layers that tend to receive much less attention and carry significant operational risk.

  • Definitional consistency. Does every part of your AI system use the same definition of a customer, a lead, a completed sale, a high-value account? When different tools or datasets carry conflicting definitions, the system will resolve the conflict invisibly — and you will not know which version it chose.
  • Policy freshness. When pricing, compliance, credit, or approval policies change inside your organisation, what is the mechanism by which an AI system that acts on those policies is updated? In most deployments, there is no mechanism. The update has to be deliberate, documented and tested.
  • Correction propagation. When a human corrects an AI output — overriding a recommendation, adjusting a classification, rejecting a draft — does that correction inform the system going forward, or does the system repeat the same mistake in the next session? Without a structured approach to memory, the answer is usually the latter.

Memory and context are not the same thing

These two terms are often used interchangeably in technology conversations, but the distinction is useful for any leader trying to govern an AI deployment.

Memory refers to what a system retains between sessions: past interactions, user preferences, prior decisions, feedback it has received. Context refers to what governs the meaning of that memory — the current definitions, the valid policies, the trusted sources, the scope within which the system is authorised to act.

A system can have excellent memory and still operate on an outdated context. It will recall exactly what was decided six months ago and apply it faithfully — even if the decision was superseded the following week. In enterprise environments this separation matters practically. The memory layer needs to optimise for recall. The context layer needs to be governed for accuracy, freshness and authority.

The question of who owns context governance — whether it sits with IT, operations, legal, or a dedicated function — is one that most organisations have not formally answered. IBM research published in 2026 found that 43 per cent of chief operations officers identify data quality as their most significant data priority, which suggests the input problem is widely recognised even if the governance response remains inconsistent.

What deliberate context governance looks like in practice

Governing the context of an AI system is not a technology project. It is closer to the discipline of maintaining a policy library or an operations manual — except that the consequences of letting it drift are faster and less visible.

In practice, organisations that manage this well tend to share a few habits. They identify, before deploying a system, which facts the system will rely on and who is responsible for keeping each of those facts current. They distinguish between information the system should retrieve in real time — live pricing, current pipeline data — and information it should receive as governed, version-controlled policy. They build review triggers: when a business rule changes, an explicit step exists to update the context the AI uses, not just the human-facing documentation.

They also establish what might be called an authorisation boundary — a clear definition of what decisions the system is permitted to reach on its own, and at what point it should surface a question rather than produce a conclusion. This is not about limiting the system's usefulness. It is about ensuring that the system's confidence is earned by the accuracy of its inputs, not merely by the fluency of its outputs.

The management question underneath the technical one

When an AI system produces a poor decision, the first instinct is often to examine the model. Was it the wrong choice? Does it need retraining? Should a different provider be evaluated? These are sometimes valid questions. But they are rarely the right first question.

The right first question is: what did this system believe when it made that decision, and was that belief accurate? If the answer to the second part is no, then the problem is governance, not capability. Adding a more capable model to an ungoverned context layer does not solve the problem. It accelerates it.

This distinction changes where leadership attention should focus. It is less about which AI tools to adopt and more about treating context — the information an AI system is permitted to act on — as a managed asset with ownership, maintenance cycles and audit checkpoints.

The organisations that will get the most reliable value from AI systems are not necessarily those with the most sophisticated models. They are those that have built the discipline to keep what their systems believe aligned with what is actually true.

Questions worth asking of any live AI deployment

The following questions are not a technical audit. They are the kind of conversation a managing director, chief operating officer or senior partner should be able to have with whoever runs their AI systems — and if the answers are not readily available, that absence is itself informative.

  • When a business policy changes, what is the process for updating the AI systems that act on it? Who initiates that update and who verifies it?
  • Which facts does each AI system treat as fixed, and how old is the most recent version of each one?
  • If the system made a consequential error tomorrow, would we be able to reconstruct what information it was working from at the time?
  • When a human overrides or corrects the system, does that correction persist — and if so, for how long and in which contexts?
  • Is there a clear boundary between what the system decides autonomously and what it refers to a human? Has that boundary been explicitly set, or has it simply accumulated by default?

Context governance is one of the less glamorous aspects of running AI systems well. It does not appear in capability demonstrations, and it rarely features in the headline metrics used to justify investment. But it is where the difference between a reliable system and a confidently wrong one is decided.

If there is a specific point in your operations where an AI system is making or informing consequential decisions — and you are not entirely certain what it believes today — that is a reasonable conversation to start. We are glad to help think it through.

Before we talk.

You do not need a solution in mind. Bring one recurring bottleneck, missed signal or decision that should work better.

Start a conversation