Artificially Confident

Artificially Confident

Practical AI, properly examined

Author: Andy

  • AI Bioresilience Is Becoming an Operating Model, Not Just a Safety Claim

    AI Bioresilience Is Becoming an Operating Model, Not Just a Safety Claim

    The most interesting AI story in biology is no longer simply whether models can accelerate science. It is whether the organisations deploying them can hold two responsibilities at once: widening legitimate research access while reducing the chance of harmful misuse.

    Google DeepMind and Isomorphic Labs’ recent bioresilience update is a useful illustration of that shift. The companies describe work with trusted partners across prevention, detection and response to biological threats, alongside safeguards intended to reduce misuse. The announcement is not proof that the hard governance questions are solved. It is evidence that the questions are now part of how frontier AI is being positioned for real-world scientific work.

    Why this matters beyond biology

    Biology is a high-stakes test case for a wider pattern in AI. The systems with the greatest potential usefulness can also create new paths to harm. That means responsible deployment cannot be reduced to a single gate at model release. It needs an operating model: which users are trusted, what they can do, how systems are evaluated, what is monitored and how concerns are escalated.

    DeepMind says its approach combines prevention, detection and response. In practical terms, it describes threat modelling, evaluations, mitigations and monitoring, as well as work with governments, researchers and biosecurity specialists. Those are familiar words in AI safety conversations, but their importance lies in the combination. A model may be capable, a policy may be well written and a partner may be credible; none of those elements is sufficient if the surrounding system cannot keep pace with changing use.

    The opportunity is real, but it is not automatic

    The update describes potential uses including helping researchers analyse sequence data, improve outbreak surveillance and accelerate the design of medical countermeasures. These are significant ambitions. They should be read as potential contributions to a broader scientific and public-health system, not as a promise that a model alone will deliver better outcomes.

    That distinction is important. In consequential domains, a technically impressive result must pass through laboratories, clinical and public-health expertise, regulation, procurement, security controls and real-world validation. AI can improve parts of the process; it does not remove the need for those institutions. The best near-term question is not “will AI solve biosecurity?” but “where can it improve the speed or quality of expert work without weakening the safeguards that make that work trustworthy?”

    Access design becomes a safety control

    One of the clearest signals in the announcement is the focus on trusted access. That approach recognises that deployment design matters. A system can have the same underlying capability but a very different risk profile depending on who can use it, which tools it can call, what information it can access and whether its outputs are independently checked.

    This is not a call for a simplistic “open versus closed” argument. Both broad and restricted access involve trade-offs. Wider availability can support research, independent scrutiny and innovation; restriction can reduce certain misuse pathways but concentrate power and make independent testing harder. Responsible access design needs to be explicit about the intended users, the evidence for the chosen safeguards and the route for revising them.

    That is an operational question as much as a technical one. An access policy that is never reviewed can become a blind spot. A review process that cannot see changes in model capability, partner use or threat information is not really a control.

    Evaluation must connect to decisions

    Model evaluations are increasingly central to frontier AI safety claims. They can test particular capabilities or failure modes under defined conditions. But an evaluation result is not self-executing. Someone must decide what it means for access, deployment, monitoring and escalation.

    That means good evaluation practice needs a decision trail: what was tested, what the limits were, who interpreted the findings, which mitigations were chosen and what would cause the judgement to be reopened. Without that trail, “we evaluated it” becomes a reassuring but incomplete statement.

    For organisations outside the frontier-model labs, the lesson is equally useful. A supplier’s safety documentation is evidence, not a substitute for local judgement. Teams should ask how their own use changes the risk: the data they connect, the actions they permit, the users they serve and the harm that could follow a failure.

    From safety statements to durable practice

    There is a temptation to treat safety language as a brand attribute. Bioresilience makes that difficult. The stakes demand specifics: named responsibilities, well-defined access conditions, independent expertise, reporting routes and a willingness to change course as capabilities or risks change.

    DeepMind’s update points to a more mature framing of the issue. It places scientific benefit and misuse prevention in the same operational picture, rather than treating them as unrelated teams or announcements. Whether that framing produces durable results will depend on the details: the robustness of evaluations, the quality of partner governance, the transparency of learning and the discipline of ongoing monitoring.

    For readers watching AI beyond the product-release cycle, that is the larger story. As models become more useful in sensitive domains, the competitive differentiator will not only be capability. It will be the credibility of the system around the capability.

    Further reading

  • From AI Pilot to Production: The Governance Handover That Matters

    From AI Pilot to Production: The Governance Handover That Matters

    An AI pilot proves that something can work. It does not prove that an organisation can responsibly rely on it.

    That distinction is where many promising AI projects stall. A team has demonstrated value: a model summarises documents, speeds up triage, improves a search workflow or drafts first-pass analysis. The next question sounds simple — “can we roll it out?” — but it is a different class of decision. A production deployment needs a durable answer to ownership, evidence, change, incidents and review.

    In other words, the work moves from experimentation to operations. That transition is a PolicyOps problem.

    A pilot is an observation; production is a commitment

    Pilots are designed to reduce uncertainty. They often run with a small group, curated data, close technical attention and a forgiving tolerance for manual workarounds. Production use introduces different conditions: more people, more data, changing inputs, dependencies on suppliers and an expectation that the service will continue to work when its original champions are busy or have moved on.

    None of that means organisations should avoid pilots. It means the evidence gathered during a pilot should be deliberately converted into a deployment decision. Without that conversion, a temporary experiment becomes a de facto service with none of the accompanying accountability.

    That risk is particularly acute with generative AI. A pilot may show that an assistant can draft useful text. It may not reveal whether staff know when to reject its output, whether sensitive material is entering the workflow, whether a supplier change will alter behaviour, or who owns the response if the system is used outside its approved context.

    The production handover should be an evidence package

    The right handover is not a single approval meeting. It is a concise package that makes the reasoning available to the people who must run and oversee the service. At a minimum, it should include:

    • Purpose and boundary: what the system is for, what it must not be used for, and which decisions remain human.
    • Accountable ownership: a named person accountable for the service, with clear responsibilities for technical operation, policy interpretation and risk acceptance.
    • Evidence: the testing, user feedback, supplier information, data assessment and known limitations that support the decision.
    • Controls: access rules, data handling, user guidance, monitoring, escalation routes and any required human review.
    • Conditions: what must remain true for the approval to remain valid.
    • Review triggers: the date and events that force a reassessment.

    This is deliberately more useful than a generic risk rating. A rating may tell a committee that a project is “medium risk”; it does not tell a future service owner what has actually been agreed. An evidence package does.

    Be precise about the change that is being approved

    “Use AI for customer service” is not an approval-ready description. It blurs together a model, a provider, data sources, an interface, an audience and a level of autonomy. Each of those can change the risk materially.

    A good production decision identifies the specific configuration: the intended user group, input and output data, model or service, integrations, human checkpoints and decisions the output may influence. That detail is not an attempt to freeze innovation. It is the baseline from which meaningful change can be measured.

    When a vendor updates a model, when a tool gains an action-taking capability, or when a team proposes using a new data source, someone should be able to compare the new reality with the approved one. If they cannot, the organisation has no reliable way to judge whether a change is minor housekeeping or a new deployment.

    Turn monitoring into a management habit

    Monitoring should also be proportionate. For a low-risk drafting assistant, useful signals might be adoption, user-reported errors, policy breaches and whether the approved data boundary is being observed. For a higher-consequence system, the organisation may need structured quality testing, disparate-impact analysis, incident metrics and regular review by a suitably independent owner.

    The important point is that the measures relate to the original decision. If a team said people would always validate outputs before use, it should be able to test whether that is happening. If a supplier gave a particular assurance, the owner should know when that assurance changes. Governance becomes real when it can detect a gap between the service that was approved and the service that is actually operating.

    Our article on decision records explains why that operational memory matters. The production handover is one of the most important records to keep: it links a policy to a real service, and it makes the underlying judgement reviewable.

    Make the exit as clear as the entry

    Production decisions also need an exit path. What happens if the provider withdraws a model, a serious issue is found, a key control fails or the organisation decides the use no longer fits its risk appetite? A simple fallback — pause the integration, revert to a manual process, preserve records, notify affected owners — can be the difference between an orderly response and an improvised one.

    This is not pessimism. It is part of treating AI as an operational capability rather than a novelty. Mature teams plan for change because change is normal.

    PolicyOps is the connective tissue

    Policies, risk assessments, supplier documents, test evidence and operational instructions often live in separate systems. That makes each document look complete while leaving the overall decision hard to understand. PolicyOps connects those elements: the policy that sets the boundary, the evidence that supports a decision, the owner who accepts it and the trigger that reopens it.

    For teams moving from pilot to production, that connection is the practical goal. The question is not merely “does this model work?” It is “can we explain why this use is acceptable, operate it within clear limits and notice when the answer changes?”

    That is the point at which an AI pilot becomes a service an organisation can stand behind.

    Further reading

    Continue the AI governance series

  • AI Makes Vulnerability Discovery Faster. That Is Not a Case for Hiding Code.

    AI Makes Vulnerability Discovery Faster. That Is Not a Case for Hiding Code.

    AI is making vulnerability discovery faster. That does not mean the sensible response is to hide every repository. The UK government’s recent guidance on AI, open code and vulnerability risk makes the more important point: the weakness is usually the weakness itself—unpatched software, unsafe configuration, exposed secrets or slow remediation—not the fact that a determined person can inspect the code.

    This distinction matters because “private by default” can sound like security while quietly reducing scrutiny, reuse and coordinated improvement. It can also distract teams from the operating capability that actually protects a service.

    AI changes speed, not the fundamentals

    Modern tools can help people analyse code, spot patterns and assemble plausible paths to a flaw more quickly. That may reduce an attacker’s uncertainty, particularly where a service is neglected. But making a repository private does not patch a known vulnerability, rotate a leaked token, harden a deployment or give an under-resourced team the ability to respond.

    The guidance is clear that public-sector organisations should meet a minimum operational standard: clear ownership, secure-by-design practice, automated hygiene and credible remediation. These are not glamorous controls, but they determine whether a vulnerability becomes an incident.

    Openness should be a deliberate default

    Open code can improve quality through reuse and review. It can also make a system easier for defenders, suppliers and peer organisations to understand. Closing code may be appropriate where there is a specific and credible route to harm, but it should be an exception with a short threat model—not a reflexive response to AI headlines.

    A useful threat model asks who might act, what access publication adds, what asset is at risk and whether the practical harm depends on code visibility. It should also be time-bound. A repository closed during active remediation should not stay closed indefinitely simply because no one revisited the decision.

    Secrets are not source code

    The most basic rule remains the most valuable: secrets do not belong in repositories, public or private. Credentials, tokens and private keys need separate controls, rotation and detection. A private repository can still be accessed, copied, exposed by a misconfiguration or shared too widely. Treating private visibility as a substitute for secret management creates a fragile boundary.

    Teams should use automated scanning before code is merged, protect branches, restrict write access, review dependencies and maintain an inventory of deployed components. These controls reduce risk regardless of whether code is visible to the public.

    Remediation is the real capability test

    AI-assisted discovery is most worrying for systems that cannot be fixed quickly. An organisation needs a known owner, an accurate view of what is running, a way to assess severity, a tested patch route and a communication process for affected users. Without those, even a privately held codebase can become a long-lived exposure.

    Measure the time between discovering a weakness and containing it. Review why a patch was delayed: was ownership unclear, did the team lack an environment, did a supplier dependency block the change, or did nobody know which service used the component? The answers are operational evidence, not just engineering detail.

    Security is a maintained condition

    The right response to AI-accelerated analysis is not secrecy as theatre. It is better hygiene, clear ownership, rapid remediation and explicit exceptions where openness genuinely adds risk. Those controls make systems safer for everyone, including the teams trying to use AI productively.

    Further reading

  • AI Governance Does Not End at Deployment

    AI Governance Does Not End at Deployment

    Most AI governance programmes put their energy into the approval decision. A use case is assessed, a policy is consulted, a supplier is reviewed and a launch is approved. Then the system enters ordinary work—and the evidence trail often ends exactly when the real risk begins.

    NIST’s 2026 work on post-deployment AI monitoring is a useful corrective. It treats monitoring as a continuing method for studying an AI system in the field: how it performs, how people use it, whether assumptions still hold and whether a real-world outcome requires action. That is not an optional dashboard. It is the operating part of assurance.

    Deployment is the start of the evidence problem

    Pre-launch evaluation matters, but it has limits. Test datasets are not the live environment. A prompt may change, a supplier may update a model, a team may expand the system’s purpose, or users may discover a workaround that changes the effective workflow. A system that passed a sensible test last quarter can still become unreliable or inappropriate today.

    The point is not to demand constant measurement of everything. It is to agree in advance what should be watched, who receives the signal and what happens when a threshold is crossed. Without those decisions, “monitoring” becomes a report that no one owns.

    Start with the decision, not the metric

    Teams often choose the metrics that are easiest to collect: usage, response time or a generic quality score. Those can be helpful, but they do not automatically show whether the system remains acceptable. Start instead with the decision the system influences and the harm that a failure could cause.

    A drafting assistant may need monitoring for confidentiality breaches, unsafe reuse and unexplained reliance. A triage system may need checks for error patterns across affected groups, escalating exception rates and whether staff can correct its recommendations. The measures should connect to a clear decision about continue, constrain, investigate, retrain or stop.

    Make a monitoring plan that somebody can operate

    A practical plan can be short. It should identify the accountable owner, intended use, current version, data boundary, monitoring signals, review cadence and escalation route. It should also say what evidence will be retained. Examples include sampled outputs, override rates, incident records, model-change notices, user feedback, evaluation results and periodic access reviews.

    That record is much more useful when it is connected to the governing policy and the original approval. Policy operations turns that connection into a working control: an owner can see what they agreed to, what has changed and what review is due next.

    Human oversight must be observable

    It is not enough to state that a person remains responsible. If an AI recommendation is routinely accepted without time, training or authority to challenge it, human oversight is nominal. Monitoring should look for signals of automation bias: unusually high acceptance, repeated overrides by the same experienced staff, short-circuited review steps or decisions made outside the intended workflow.

    A good escalation route does not punish people for raising uncertainty. It makes the next action straightforward: pause a capability, remove a data source, add a check, retrain users or refer the case to a named owner. The organisation needs to be able to act on the signal, not merely record it.

    Review material changes deliberately

    A new model version, a different data source, a new supplier term or a broader user group can all alter the risk profile. Treat these as events that trigger review rather than as routine maintenance. The question is not whether any change occurred; it is whether the change affects the purpose, performance, data, authority or potential impact of the system.

    This discipline keeps assurance current without turning every small improvement into a committee meeting. Low-impact changes can follow a light route. Material changes should reopen the evidence, controls and approval decision.

    Monitoring creates a more honest form of confidence

    No organisation can prove that an AI system will never fail. It can show that it knows what good performance looks like, has a way to detect when reality diverges from that expectation and can respond. That is a much stronger claim than “approved once.”

    Further reading

  • AI Agents Need Identity and Authority, Not Just Better Prompts

    AI Agents Need Identity and Authority, Not Just Better Prompts

    AI agents change the security question from “what did the model say?” to “what was the system allowed to do?” Once an agent can call tools, access records or trigger workflows, its identity and authority become part of the security boundary.

    NIST is treating agent identity as infrastructure

    In 2026, NIST’s Center for AI Standards and Innovation published work on security considerations for AI agents and a concept paper on identity and authorization for software agents. The direction is significant: agent security is not only about model behaviour or prompt injection. It is also about connecting a delegated action to a real user, a defined purpose and an auditable permission.

    That is more useful than treating an agent as an unusually clever service account. A service account with broad access is already difficult to govern. An agent adds planning, interpretation and tool selection, increasing the ways a legitimate instruction can produce an unsafe result.

    Delegation needs a boundary

    An agent should not inherit every permission available to the person who started it. Delegation should be explicit and narrow: which user initiated the task, which agent is acting, which tools are available, which resources are in scope, how long authority lasts and what actions require confirmation.

    If an agent calls a second agent or an external service, downstream action should not lose the original context. Systems need to preserve the relationship between the human request, agent plan, tool call, result and final action.

    Logs must explain actions, not just errors

    Many systems log authentication and application errors but not the decision path that matters most. An agent audit record should show the triggering request, relevant policy, tools considered, tool called, data returned, approval checkpoint and resulting change.

    This does not mean storing every internal model token. It means enough structured evidence to reconstruct responsibility. A reviewer should tell whether the agent acted within authority, whether a human approved the right step and whether the system had enough information to make the action safe.

    Least privilege is necessary but not sufficient

    An agent that only needs to read a ticket should not be able to close it, change an entitlement or export a customer list. But permissions alone do not solve context confusion. A read permission can expose sensitive information, and a valid write permission can still be inappropriate for a particular case.

    Controls should combine permissions with purpose, data classification, transaction limits and separation of duties. High-impact actions may need a second person, restricted workflow or reversible staging step. The control should depend on the consequence of failure, not whether the interface says “assistant” or “agent”.

    Make the failure path part of the design

    Agent systems encounter ambiguous instructions, malicious content, unavailable tools and conflicting records. A safe design makes refusal and escalation normal outcomes. The agent should stop, explain what is missing and route the task to an owner without trying increasingly broad actions.

    Testing should include indirect instructions in retrieved documents, poisoned tool responses, stale permissions, duplicate requests, partial failures and attempts to cross a data boundary. NIST’s work points toward shared standards, but each organisation still has to test the concrete agent and workflow it operates.

    What teams can implement now

    • Give each production agent a distinct identity and owner.
    • Record the human principal, purpose and expiry for delegated work.
    • Use narrowly scoped tools with separate read and write permissions.
    • Require confirmation or second-party approval for consequential actions.
    • Log tool calls, approvals, data boundaries and resulting changes.
    • Review dormant agents and revoke access when ownership changes.

    The shift is from model trust to action accountability

    Capability is not authority. The safer pattern is to treat an agent as a delegated operator whose identity, permissions and actions remain visible throughout the workflow. That gives organisations a practical basis for trust: not confidence that the agent will always behave well, but evidence that failures can be bounded, detected and corrected.

    Further reading

  • The EU AI Act Is Becoming an Operational Readiness Test

    The EU AI Act Is Becoming an Operational Readiness Test

    The EU AI Act is moving from a compliance timetable into an operating problem. For many organisations, the important question is no longer whether the regulation exists, but whether they can show what an AI system does, who owns it, what evidence supports its use, and how it is monitored after launch.

    August 2026 is a planning deadline, not a paperwork deadline

    The European Commission says the Act entered into force on 1 August 2024 and is generally due to become fully applicable on 2 August 2026, subject to specific transition periods. The AI Act Service Desk timeline identifies obligations that phase in at different dates. That structure makes a simple “we will comply in August” plan risky: different systems, providers and duties may arrive at different points.

    The practical response is to build an evidence-backed inventory now. Identify the system, provider, model or service dependency, business purpose, affected users, data boundary, decision-maker, risk classification and current status. Record uncertainty too. A system should not become “low risk” merely because no one has yet taken responsibility for classifying it.

    Turn each system into an accountable object

    A policy describes an organisation’s intent. Operational readiness requires a system record that can be inspected and updated. It should answer: what decision or task does the system support; who owns and operates it; what data enters and leaves; what tests and approvals support deployment; and what would cause it to be paused, reviewed or retired?

    This is where policy operations matters. Governance is stronger when policy requirements are linked to owned controls, evidence and review events, rather than stored as disconnected documents.

    Evidence should follow the lifecycle

    Teams often gather evidence for an approval and then stop. That creates a false sense of readiness. A deployed AI system changes as its model, prompt, data, users and surrounding workflow change. Evidence should therefore cover design, procurement, testing, approval, deployment, monitoring, incident response and retirement.

    Useful evidence might include a signed use-case decision, supplier assessment, evaluation results, data-flow diagram, human-oversight procedure, training record, monitoring results and dated review decision. The key property is traceability: a reviewer should understand why a control exists, what it covers and whether it is current.

    Do not confuse a deadline with assurance

    Implementation guidance can clarify obligations, but it cannot prove that a particular deployment is safe or appropriate. Organisations still have to judge accuracy, discrimination, privacy, security, resilience and the consequences of error. Those judgments should be proportionate, but they should not disappear into a generic risk score.

    Ask what happens when the system is wrong, unavailable, manipulated or used outside its intended context. Can a person detect the failure, intervene in time and explain what happened afterwards? For consequential uses, the answer should be supported by tested procedures rather than an assumption that a human is “in the loop”.

    A practical readiness sequence

    First, identify active and planned AI use cases. Second, assign an accountable owner. Third, map the data, suppliers, users and decisions. Fourth, identify missing evidence and controls. Fifth, schedule reviews around material changes, incidents and regulatory milestones.

    This creates a defensible record of preparation. It distinguishes a policy that has been published from a control that has been implemented, tested and reviewed. That distinction will matter well beyond August 2026.

    The useful question to ask now

    Instead of asking whether the organisation has an AI policy, ask whether it can explain the current state of every material AI system in one place. If it cannot identify the owner, evidence, boundaries and next review, the remaining work is operational—not cosmetic.

    Further reading

  • AI Governance Needs Decision Logs, Not Just Policies

    AI Governance Needs Decision Logs, Not Just Policies

    Most AI governance failures are not caused by a missing policy. They happen when nobody can reconstruct how a policy was interpreted, who made a decision, what evidence they used, or whether the decision was ever revisited.

    That is why the most useful unit of PolicyOps is often not the policy document. It is the decision record: a short, durable account of what was decided, by whom, under which rule, with what evidence, and when it must next be reviewed.

    As organisations put generative AI into real work, this distinction becomes practical. A policy can say that sensitive information must not be entered into a public model, that people must remain accountable for consequential decisions, or that suppliers must be assessed. Those are important constraints. But a policy alone cannot answer the questions that arrive later: was this particular use case approved? Was the model’s scope understood? Which data path was accepted? Who agreed the residual risk? What changed after deployment?

    A credible operating model needs a memory.

    Policies set direction. Decision records make it operational.

    Policies are deliberately general. They establish principles, boundaries and responsibilities that should endure beyond one product or project. Decisions are local: a team wants to use a specific model, for a defined purpose, with particular data, controls and owners. Treating the two as the same thing creates a familiar gap. The policy exists, the project moves quickly, and the evidence of interpretation is scattered between a ticket, a meeting, a procurement folder and someone’s memory.

    A lightweight decision record closes that gap without turning governance into a ceremony. It should capture:

    • the decision being made and the business context;
    • the policy, standard or legal obligation that applies;
    • the evidence considered, including supplier material and testing;
    • the accountable owner and any reviewers or approvers;
    • the safeguards, assumptions and residual risks accepted; and
    • a review trigger: a date, material model change, incident, new data source or change in use.

    That is not bureaucracy for its own sake. It makes a decision legible to the next person who has to operate, challenge, audit or improve it. It also stops a generic approval from silently becoming permission for a different system, dataset or purpose six months later.

    Why AI makes the evidence problem sharper

    AI systems shift in ways that ordinary software procurement often does not. A model provider may change behaviour, an integration may begin carrying a new category of data, or users may find a valuable use that was never part of the original assessment. The risk is not just a model producing an odd answer. It is an organisation continuing to rely on an old judgement after the facts supporting that judgement have changed.

    The NIST AI Risk Management Framework frames AI risk management as an ongoing activity across design, development, use and evaluation. Its companion work on generative AI similarly treats risks as contextual rather than as a one-time checklist. That is a useful corrective: governance should be able to show not only that a control was named, but how it was applied to a particular use.

    Decision records are the bridge between that principle and day-to-day work. A supplier questionnaire may show what a vendor said at a point in time. A testing report may show what was observed. A decision record connects those inputs to an accountable judgement: this use was accepted for this purpose, subject to these conditions, until this review event.

    Make review triggers explicit

    The most overlooked field in a decision record is the trigger to reopen it. A date is useful, but it is not enough. Good governance is also event-driven. A decision should return for review when, for example:

    • the model, provider, hosting location or core capability changes;
    • a new data type enters the workflow;
    • the system moves from assistance to automation, or affects a consequential decision;
    • monitoring identifies a meaningful performance, bias, security or privacy concern; or
    • a policy, regulatory expectation or internal risk appetite changes.

    These triggers turn a static assessment into a living control. They also create a more honest conversation with delivery teams: approval is not a permanent green light. It is a bounded judgement, made under stated conditions.

    This is especially valuable in public-sector and other high-consequence settings, where explainability needs to include the organisation’s own choices. Our recent look at UK police use of AI made the point from a different angle: capability can expand faster than the operating model that gives it legitimacy. A durable trail of decisions does not make a difficult use case acceptable by itself, but it makes challenge, oversight and correction possible.

    Keep the record proportionate

    Not every decision needs a board paper. The record should be proportionate to the potential impact. A low-risk internal drafting assistant might require a simple owner, approved data boundary, supplier assessment and annual review. A system that influences eligibility, enforcement, safety, employment or access to services needs much more: clear authority, stronger evidence, testing, affected-party considerations and a route to challenge.

    What matters is consistency. If every team invents its own wording and storage location, leaders cannot see the portfolio of decisions they are carrying. They cannot spot repeated dependencies on one supplier, a control that is repeatedly waived, or a cluster of projects working from the same outdated assumption.

    That is the practical promise of PolicyOps: policies, evidence, decisions, ownership and review should be connected rather than treated as separate administrative tasks. The aim is not to centralise every judgement. It is to make authorised judgement traceable, reviewable and easier to improve.

    Start with one decision that matters

    Teams do not need to redesign their entire governance estate to begin. Pick one live AI use case that has crossed from experiment into recurring work. Identify the accountable owner. Link the relevant policy. Write down the evidence and conditions that support the current decision. Then choose the event that would force the team to revisit it.

    That simple exercise usually reveals the real gaps: an unclear owner, an assumption about data that has never been checked, no agreed route for incidents, or no way to tell whether the supplier has materially changed the service. Those are not paperwork problems. They are operating-model problems.

    AI governance becomes credible when it can answer a modest but demanding question: why are we allowed to do this, and how will we know when that answer is no longer good enough? A well-kept decision record is where that answer lives.

    Further reading

    Continue the AI governance series

  • UK Police AI Is Expanding Faster Than Its Operating Model

    UK Police AI Is Expanding Faster Than Its Operating Model

    Britain is about to put more AI into policing. The government has announced PoliceAI, a national centre intended to speed up adoption, with pilots for triaging and summarising digital evidence, more live facial-recognition capacity and a planned public register of tools in use. The case for some of this is straightforward: investigations now generate vast quantities of digital material, and officers should not have to spend their working lives manually redacting video or searching duplicated files.

    But the most revealing recent UK police AI story is not a successful pilot. It is the parliamentary inquiry into the decision to exclude Maccabi Tel Aviv supporters from an Aston Villa fixture last year. The Home Affairs Committee found that West Midlands Police had relied on inaccurate and unverified information, including material generated through Microsoft Copilot, when building part of its intelligence picture. That information was used in a decision with real consequences for people, public trust and community relations.

    This should not be turned into the lazy conclusion that every police officer is incapable of using AI, or that every AI-assisted task is inherently unsafe. It is more serious than that. It shows what happens when an organisation introduces a persuasive new tool without an equally clear operating model for checking, escalating, recording and owning its use.

    The problem is not simply accuracy

    Generative AI gets things wrong. That is not news. It can produce a confident summary that contains invented detail, merge unrelated events or repeat a weak source in fluent language. In a low-stakes setting, the answer may be an embarrassing error. In policing, it can alter how risk is framed, who is treated as a threat and whether a decision survives scrutiny.

    The Committee’s report is striking because it describes more than a bad output. It describes a broken decision route. Material that supported a pre-existing narrative was accepted; contradictory evidence from authoritative sources was not given enough weight. A fictitious fixture and other claims entered the account. The use of AI was itself not properly surfaced or understood at senior level in time for accurate evidence to Parliament. That is not a prompt-writing failure. It is a governance failure.

    Every public authority already knows, in principle, that intelligence should be assessed, sourced and challenged. AI does not remove those obligations. It increases the need to apply them, because it makes it cheap to generate a lot of plausible-looking material very quickly. The danger is not that a model replaces professional judgement overnight. The danger is that uncertain material quietly acquires the status of professional judgement as it moves through a briefing, a meeting and a decision.

    The national response acknowledges the gap

    There is a tension at the centre of the current policy moment. Ministers are rightly focused on the operational upside: faster handling of digital evidence, less repetitive work and more capacity for officers to investigate. At the same time, the new programme promises a public registry, independent testing for accuracy and bias, and governance support. Those are welcome commitments. They also reveal that the common baseline has not yet been fully built.

    That matters because police forces do not deploy AI into a neutral environment. They use it alongside existing powers, intelligence processes, data-protection duties and decisions that can affect liberty, safety and community confidence. The higher the consequence, the less adequate it is to say that a human remains in the loop. The real question is what that human is expected to do, what evidence they can see, and whether they have enough time and authority to challenge the output.

    Recent research from Northumbria University makes a similar point. Its mapping of probabilistic AI across the criminal justice system found adoption moving faster than the safeguards intended to govern it. The useful takeaway is not a call to freeze every experiment. It is that scale without an operating discipline creates a patchwork: one team may have robust checking and records, while another treats an AI-generated answer as a useful shortcut and moves on.

    What “knowing how to use it” actually means

    Knowing how to use AI in policing is not just knowing where the button is. It means being able to answer a small set of practical questions before an output influences a real decision.

    • What is the tool doing? Is it retrieving material, summarising it, ranking risk, generating prose or identifying a person? These are not interchangeable activities and should not have the same controls.
    • What is the source? Can an officer or decision-maker trace an assertion back to the original intelligence, record or evidence? A fluent answer without provenance is not a reliable briefing.
    • What must be checked? A policy needs to make clear which claims need independent verification, who performs it and when that check is recorded.
    • Who owns the decision? “Human in the loop” is too vague. Someone needs named responsibility for accepting, rejecting or escalating a material output.
    • What happens when the tool is wrong? There should be a route to correct the record, notify people affected where appropriate, learn from the failure and stop repeated use of a faulty pattern.

    None of this is exotic. It is the discipline public institutions already use for other kinds of evidence and operational decision-making. The difference is that AI can make it easier to skip the visible parts of that discipline. A system that gives an answer in seconds can make the underlying uncertainty feel smaller than it is.

    Transparency is part of operational quality

    The promised police AI register is therefore more than a communications exercise. A clear public account of which tools are used, for what purpose, with what data and under what assurance helps create the pressure for better internal practice. It gives communities, oversight bodies and frontline staff something concrete to examine. It also forces a basic distinction that is too often blurred: a tool that organises case files is not the same as a tool that influences a stop, a watchlist, an investigation or a public-order decision.

    Transparency alone is not enough. A register can become a list of product names without answering whether a particular force has trained people, completed an impact assessment, tested for error and bias, or established a meaningful challenge process. But secrecy is worse. If a public body cannot explain the role an AI system played after the fact, it has probably not created a sufficiently accountable route for using it beforehand.

    The lesson from West Midlands is not to stop innovating

    There will be a temptation to treat the Maccabi case as an awkward exception: an individual mistake, a moment of poor judgement, then move on. That would miss the value of the warning. High-consequence use exposes weaknesses that can sit unnoticed in lower-stakes workflows. If a force cannot show how AI-generated material was checked before it affected a sensitive public decision, then the organisation does not yet have the controls required to scale that use safely.

    Conversely, a better operating model would not make police work slower or more bureaucratic for the sake of it. It would separate routine assistance from consequential advice. It would make verified sources easy to access, mark AI-derived material clearly, require confirmation at defined points, and preserve a record of the judgement made. Good controls reduce the time wasted later on correction, inquiry and loss of confidence.

    That is the real PolicyOps question in public services: not whether an organisation has a policy saying “use AI responsibly,” but whether that policy changes what people do on a busy day. Can an officer see the boundary? Can a supervisor challenge the output? Can an affected community understand what happened? Can an inspector reconstruct the route from input to decision?

    Where CopPlan fits

    This is exactly the terrain that CopPlan is designed for: helping investigative policing turn work from disparate operational systems, case files, statements, body-worn video and guidance into a clearer supervised workflow. That is a more useful ambition than an AI tool that simply produces an answer. In a policing context, the value is in helping officers and supervisors see the case, the next action and the supporting material without making the system the decision-maker.

    The distinction matters. Responsible AI for investigations should reduce administrative drag and improve visibility, while leaving authority with accountable people. It should help users locate the source, identify gaps, prioritise the work and understand why a recommendation appears. It should not turn a probabilistic output into unexamined intelligence. Building those constraints into the workflow is how a platform earns trust in a setting where errors can affect real people.

    Move quickly, but make the route visible

    UK policing should use technology that genuinely helps it investigate fairly and effectively. The case for better tools to handle digital evidence is strong. But faster adoption cannot be the only measure of success. The standard has to be whether an AI-assisted decision is more accurate, more accountable and easier to explain than the process it replaces.

    PoliceAI could be an opportunity to establish that standard nationally rather than leaving each force to invent it under pressure. The early commitments on independent testing, transparency and governance point in the right direction. Now they need to become routine practice, not future aspirations. In a public institution with coercive powers, the answer is never simply “the computer suggested it.”

    Sources: House of Commons Home Affairs Committee on the Maccabi Tel Aviv fan-ban inquiry; GOV.UK: PoliceAI announcement; Northumbria University research on safeguards and probabilistic AI.

  • Health Data in AI Assistants Changes the Question

    Health Data in AI Assistants Changes the Question

    AI has been moving steadily closer to personal data. This week, that move became more concrete: OpenAI announced Health in ChatGPT, a U.S. rollout that lets people choose to connect Apple Health and supported medical records to their conversations.

    The interesting part is not that an assistant can discuss health. People already ask AI health questions. The important change is that a system can now work with more of the context behind those questions: prior records, activity, medications and results, when a person chooses to connect them.

    Context is useful. It is also sensitive.

    More context can make an answer more useful. It can help somebody prepare questions for an appointment, notice a change over time or understand a medical note in plainer language. But health data is not just another personalisation signal. The consequences of getting privacy, permissions or accuracy wrong are much higher.

    OpenAI says connected health information is not used to train foundation models or target ads, and that users control when it can be used. It also says the product is not a replacement for professional care. Those boundaries matter—but they are only meaningful when people can understand and exercise them in the product.

    The new standard is understandable control

    • Permission should be specific: people need to know what is being connected and when it will be used.
    • Data should be easy to withdraw: disconnecting should be as clear as connecting.
    • Limits should be visible: an AI explanation is not a diagnosis or a clinical decision.
    • Important claims should remain checkable: users need a route back to the original record and to professional advice.

    This is a useful test case for AI more broadly. The closer an assistant gets to consequential personal context, the less acceptable vague privacy language becomes. Trust will depend on granular permissions, clear defaults and a real ability to change one’s mind.

    Source: OpenAI’s Health in ChatGPT announcement, published 23 July 2026.

  • PolicyOps Starts Where the Policy PDF Ends

    PolicyOps Starts Where the Policy PDF Ends

    Most policy programmes end at the point where the document is published. That is understandable: getting a policy agreed is difficult work. But the PDF is not the operating model. It is the start of one.

    The real test arrives when somebody needs an answer quickly. Can a team use a new AI tool? Which data can go into it? Who can approve an exception? What evidence should be kept? A folder full of policies may contain the answer, but it rarely provides a usable route to it.

    From document to decision

    PolicyOps starts with a simple shift: treat the policy as a source for decisions, not a final destination. That means making the approved version findable, giving people the relevant context, and preserving the path from question to outcome.

    Good policy operations do not turn every question into a compliance theatre exercise. They reduce the friction around routine decisions while making the important ones more visible. If the answer is straightforward, people should be able to see the source and move. If it is uncertain, the escalation path should be clear.

    What a usable policy system needs

    • A source of truth: current, approved material rather than a remembered summary.
    • Context: the relevant clause, related policies and limits behind a short answer.
    • Ownership: a visible route for exceptions and decisions that need judgement.
    • Evidence: enough of a record to explain what happened later.

    AI can make this experience faster, but it should not make authority disappear. A useful assistant can retrieve, summarise and connect material. It cannot quietly become the owner of the decision.

    The practical advantage

    When policy guidance is genuinely operational, people spend less time asking where the rule lives and more time applying it well. That is especially valuable in AI governance, where the questions evolve faster than most policy libraries do.

    PolicyOps is a useful example of the approach: governed answers from approved sources, with evidence and human ownership kept in view. The point is not a cleverer policy PDF. It is a better decision route.