Most AI governance failures are not caused by a missing policy. They happen when nobody can reconstruct how a policy was interpreted, who made a decision, what evidence they used, or whether the decision was ever revisited.
That is why the most useful unit of PolicyOps is often not the policy document. It is the decision record: a short, durable account of what was decided, by whom, under which rule, with what evidence, and when it must next be reviewed.
As organisations put generative AI into real work, this distinction becomes practical. A policy can say that sensitive information must not be entered into a public model, that people must remain accountable for consequential decisions, or that suppliers must be assessed. Those are important constraints. But a policy alone cannot answer the questions that arrive later: was this particular use case approved? Was the model’s scope understood? Which data path was accepted? Who agreed the residual risk? What changed after deployment?
A credible operating model needs a memory.
Policies set direction. Decision records make it operational.
Policies are deliberately general. They establish principles, boundaries and responsibilities that should endure beyond one product or project. Decisions are local: a team wants to use a specific model, for a defined purpose, with particular data, controls and owners. Treating the two as the same thing creates a familiar gap. The policy exists, the project moves quickly, and the evidence of interpretation is scattered between a ticket, a meeting, a procurement folder and someone’s memory.
A lightweight decision record closes that gap without turning governance into a ceremony. It should capture:
- the decision being made and the business context;
- the policy, standard or legal obligation that applies;
- the evidence considered, including supplier material and testing;
- the accountable owner and any reviewers or approvers;
- the safeguards, assumptions and residual risks accepted; and
- a review trigger: a date, material model change, incident, new data source or change in use.
That is not bureaucracy for its own sake. It makes a decision legible to the next person who has to operate, challenge, audit or improve it. It also stops a generic approval from silently becoming permission for a different system, dataset or purpose six months later.
Why AI makes the evidence problem sharper
AI systems shift in ways that ordinary software procurement often does not. A model provider may change behaviour, an integration may begin carrying a new category of data, or users may find a valuable use that was never part of the original assessment. The risk is not just a model producing an odd answer. It is an organisation continuing to rely on an old judgement after the facts supporting that judgement have changed.
The NIST AI Risk Management Framework frames AI risk management as an ongoing activity across design, development, use and evaluation. Its companion work on generative AI similarly treats risks as contextual rather than as a one-time checklist. That is a useful corrective: governance should be able to show not only that a control was named, but how it was applied to a particular use.
Decision records are the bridge between that principle and day-to-day work. A supplier questionnaire may show what a vendor said at a point in time. A testing report may show what was observed. A decision record connects those inputs to an accountable judgement: this use was accepted for this purpose, subject to these conditions, until this review event.
Make review triggers explicit
The most overlooked field in a decision record is the trigger to reopen it. A date is useful, but it is not enough. Good governance is also event-driven. A decision should return for review when, for example:
- the model, provider, hosting location or core capability changes;
- a new data type enters the workflow;
- the system moves from assistance to automation, or affects a consequential decision;
- monitoring identifies a meaningful performance, bias, security or privacy concern; or
- a policy, regulatory expectation or internal risk appetite changes.
These triggers turn a static assessment into a living control. They also create a more honest conversation with delivery teams: approval is not a permanent green light. It is a bounded judgement, made under stated conditions.
This is especially valuable in public-sector and other high-consequence settings, where explainability needs to include the organisation’s own choices. Our recent look at UK police use of AI made the point from a different angle: capability can expand faster than the operating model that gives it legitimacy. A durable trail of decisions does not make a difficult use case acceptable by itself, but it makes challenge, oversight and correction possible.
Keep the record proportionate
Not every decision needs a board paper. The record should be proportionate to the potential impact. A low-risk internal drafting assistant might require a simple owner, approved data boundary, supplier assessment and annual review. A system that influences eligibility, enforcement, safety, employment or access to services needs much more: clear authority, stronger evidence, testing, affected-party considerations and a route to challenge.
What matters is consistency. If every team invents its own wording and storage location, leaders cannot see the portfolio of decisions they are carrying. They cannot spot repeated dependencies on one supplier, a control that is repeatedly waived, or a cluster of projects working from the same outdated assumption.
That is the practical promise of PolicyOps: policies, evidence, decisions, ownership and review should be connected rather than treated as separate administrative tasks. The aim is not to centralise every judgement. It is to make authorised judgement traceable, reviewable and easier to improve.
Start with one decision that matters
Teams do not need to redesign their entire governance estate to begin. Pick one live AI use case that has crossed from experiment into recurring work. Identify the accountable owner. Link the relevant policy. Write down the evidence and conditions that support the current decision. Then choose the event that would force the team to revisit it.
That simple exercise usually reveals the real gaps: an unclear owner, an assumption about data that has never been checked, no agreed route for incidents, or no way to tell whether the supplier has materially changed the service. Those are not paperwork problems. They are operating-model problems.
AI governance becomes credible when it can answer a modest but demanding question: why are we allowed to do this, and how will we know when that answer is no longer good enough? A well-kept decision record is where that answer lives.
Further reading
- NIST AI Risk Management Framework
- More Artificially Confident articles on AI governance
- How Artificially Confident approaches assisted editorial work

