Most AI governance programmes put their energy into the approval decision. A use case is assessed, a policy is consulted, a supplier is reviewed and a launch is approved. Then the system enters ordinary work—and the evidence trail often ends exactly when the real risk begins.
NIST’s 2026 work on post-deployment AI monitoring is a useful corrective. It treats monitoring as a continuing method for studying an AI system in the field: how it performs, how people use it, whether assumptions still hold and whether a real-world outcome requires action. That is not an optional dashboard. It is the operating part of assurance.
Deployment is the start of the evidence problem
Pre-launch evaluation matters, but it has limits. Test datasets are not the live environment. A prompt may change, a supplier may update a model, a team may expand the system’s purpose, or users may discover a workaround that changes the effective workflow. A system that passed a sensible test last quarter can still become unreliable or inappropriate today.
The point is not to demand constant measurement of everything. It is to agree in advance what should be watched, who receives the signal and what happens when a threshold is crossed. Without those decisions, “monitoring” becomes a report that no one owns.
Start with the decision, not the metric
Teams often choose the metrics that are easiest to collect: usage, response time or a generic quality score. Those can be helpful, but they do not automatically show whether the system remains acceptable. Start instead with the decision the system influences and the harm that a failure could cause.
A drafting assistant may need monitoring for confidentiality breaches, unsafe reuse and unexplained reliance. A triage system may need checks for error patterns across affected groups, escalating exception rates and whether staff can correct its recommendations. The measures should connect to a clear decision about continue, constrain, investigate, retrain or stop.
Make a monitoring plan that somebody can operate
A practical plan can be short. It should identify the accountable owner, intended use, current version, data boundary, monitoring signals, review cadence and escalation route. It should also say what evidence will be retained. Examples include sampled outputs, override rates, incident records, model-change notices, user feedback, evaluation results and periodic access reviews.
That record is much more useful when it is connected to the governing policy and the original approval. Policy operations turns that connection into a working control: an owner can see what they agreed to, what has changed and what review is due next.
Human oversight must be observable
It is not enough to state that a person remains responsible. If an AI recommendation is routinely accepted without time, training or authority to challenge it, human oversight is nominal. Monitoring should look for signals of automation bias: unusually high acceptance, repeated overrides by the same experienced staff, short-circuited review steps or decisions made outside the intended workflow.
A good escalation route does not punish people for raising uncertainty. It makes the next action straightforward: pause a capability, remove a data source, add a check, retrain users or refer the case to a named owner. The organisation needs to be able to act on the signal, not merely record it.
Review material changes deliberately
A new model version, a different data source, a new supplier term or a broader user group can all alter the risk profile. Treat these as events that trigger review rather than as routine maintenance. The question is not whether any change occurred; it is whether the change affects the purpose, performance, data, authority or potential impact of the system.
This discipline keeps assurance current without turning every small improvement into a committee meeting. Low-impact changes can follow a light route. Material changes should reopen the evidence, controls and approval decision.
Monitoring creates a more honest form of confidence
No organisation can prove that an AI system will never fail. It can show that it knows what good performance looks like, has a way to detect when reality diverges from that expectation and can respond. That is a much stronger claim than “approved once.”

