Artificially Confident

Artificially Confident

Practical AI, properly examined

Category: AI Governance

Practical AI governance: how organisations turn policies, accountability and evidence into better decisions.

  • AI Governance Does Not End at Deployment

    AI Governance Does Not End at Deployment

    Most AI governance programmes put their energy into the approval decision. A use case is assessed, a policy is consulted, a supplier is reviewed and a launch is approved. Then the system enters ordinary work—and the evidence trail often ends exactly when the real risk begins.

    NIST’s 2026 work on post-deployment AI monitoring is a useful corrective. It treats monitoring as a continuing method for studying an AI system in the field: how it performs, how people use it, whether assumptions still hold and whether a real-world outcome requires action. That is not an optional dashboard. It is the operating part of assurance.

    Deployment is the start of the evidence problem

    Pre-launch evaluation matters, but it has limits. Test datasets are not the live environment. A prompt may change, a supplier may update a model, a team may expand the system’s purpose, or users may discover a workaround that changes the effective workflow. A system that passed a sensible test last quarter can still become unreliable or inappropriate today.

    The point is not to demand constant measurement of everything. It is to agree in advance what should be watched, who receives the signal and what happens when a threshold is crossed. Without those decisions, “monitoring” becomes a report that no one owns.

    Start with the decision, not the metric

    Teams often choose the metrics that are easiest to collect: usage, response time or a generic quality score. Those can be helpful, but they do not automatically show whether the system remains acceptable. Start instead with the decision the system influences and the harm that a failure could cause.

    A drafting assistant may need monitoring for confidentiality breaches, unsafe reuse and unexplained reliance. A triage system may need checks for error patterns across affected groups, escalating exception rates and whether staff can correct its recommendations. The measures should connect to a clear decision about continue, constrain, investigate, retrain or stop.

    Make a monitoring plan that somebody can operate

    A practical plan can be short. It should identify the accountable owner, intended use, current version, data boundary, monitoring signals, review cadence and escalation route. It should also say what evidence will be retained. Examples include sampled outputs, override rates, incident records, model-change notices, user feedback, evaluation results and periodic access reviews.

    That record is much more useful when it is connected to the governing policy and the original approval. Policy operations turns that connection into a working control: an owner can see what they agreed to, what has changed and what review is due next.

    Human oversight must be observable

    It is not enough to state that a person remains responsible. If an AI recommendation is routinely accepted without time, training or authority to challenge it, human oversight is nominal. Monitoring should look for signals of automation bias: unusually high acceptance, repeated overrides by the same experienced staff, short-circuited review steps or decisions made outside the intended workflow.

    A good escalation route does not punish people for raising uncertainty. It makes the next action straightforward: pause a capability, remove a data source, add a check, retrain users or refer the case to a named owner. The organisation needs to be able to act on the signal, not merely record it.

    Review material changes deliberately

    A new model version, a different data source, a new supplier term or a broader user group can all alter the risk profile. Treat these as events that trigger review rather than as routine maintenance. The question is not whether any change occurred; it is whether the change affects the purpose, performance, data, authority or potential impact of the system.

    This discipline keeps assurance current without turning every small improvement into a committee meeting. Low-impact changes can follow a light route. Material changes should reopen the evidence, controls and approval decision.

    Monitoring creates a more honest form of confidence

    No organisation can prove that an AI system will never fail. It can show that it knows what good performance looks like, has a way to detect when reality diverges from that expectation and can respond. That is a much stronger claim than “approved once.”

    Further reading

  • The EU AI Act Is Becoming an Operational Readiness Test

    The EU AI Act Is Becoming an Operational Readiness Test

    The EU AI Act is moving from a compliance timetable into an operating problem. For many organisations, the important question is no longer whether the regulation exists, but whether they can show what an AI system does, who owns it, what evidence supports its use, and how it is monitored after launch.

    August 2026 is a planning deadline, not a paperwork deadline

    The European Commission says the Act entered into force on 1 August 2024 and is generally due to become fully applicable on 2 August 2026, subject to specific transition periods. The AI Act Service Desk timeline identifies obligations that phase in at different dates. That structure makes a simple “we will comply in August” plan risky: different systems, providers and duties may arrive at different points.

    The practical response is to build an evidence-backed inventory now. Identify the system, provider, model or service dependency, business purpose, affected users, data boundary, decision-maker, risk classification and current status. Record uncertainty too. A system should not become “low risk” merely because no one has yet taken responsibility for classifying it.

    Turn each system into an accountable object

    A policy describes an organisation’s intent. Operational readiness requires a system record that can be inspected and updated. It should answer: what decision or task does the system support; who owns and operates it; what data enters and leaves; what tests and approvals support deployment; and what would cause it to be paused, reviewed or retired?

    This is where policy operations matters. Governance is stronger when policy requirements are linked to owned controls, evidence and review events, rather than stored as disconnected documents.

    Evidence should follow the lifecycle

    Teams often gather evidence for an approval and then stop. That creates a false sense of readiness. A deployed AI system changes as its model, prompt, data, users and surrounding workflow change. Evidence should therefore cover design, procurement, testing, approval, deployment, monitoring, incident response and retirement.

    Useful evidence might include a signed use-case decision, supplier assessment, evaluation results, data-flow diagram, human-oversight procedure, training record, monitoring results and dated review decision. The key property is traceability: a reviewer should understand why a control exists, what it covers and whether it is current.

    Do not confuse a deadline with assurance

    Implementation guidance can clarify obligations, but it cannot prove that a particular deployment is safe or appropriate. Organisations still have to judge accuracy, discrimination, privacy, security, resilience and the consequences of error. Those judgments should be proportionate, but they should not disappear into a generic risk score.

    Ask what happens when the system is wrong, unavailable, manipulated or used outside its intended context. Can a person detect the failure, intervene in time and explain what happened afterwards? For consequential uses, the answer should be supported by tested procedures rather than an assumption that a human is “in the loop”.

    A practical readiness sequence

    First, identify active and planned AI use cases. Second, assign an accountable owner. Third, map the data, suppliers, users and decisions. Fourth, identify missing evidence and controls. Fifth, schedule reviews around material changes, incidents and regulatory milestones.

    This creates a defensible record of preparation. It distinguishes a policy that has been published from a control that has been implemented, tested and reviewed. That distinction will matter well beyond August 2026.

    The useful question to ask now

    Instead of asking whether the organisation has an AI policy, ask whether it can explain the current state of every material AI system in one place. If it cannot identify the owner, evidence, boundaries and next review, the remaining work is operational—not cosmetic.

    Further reading

  • AI Governance Needs Decision Logs, Not Just Policies

    AI Governance Needs Decision Logs, Not Just Policies

    Most AI governance failures are not caused by a missing policy. They happen when nobody can reconstruct how a policy was interpreted, who made a decision, what evidence they used, or whether the decision was ever revisited.

    That is why the most useful unit of PolicyOps is often not the policy document. It is the decision record: a short, durable account of what was decided, by whom, under which rule, with what evidence, and when it must next be reviewed.

    As organisations put generative AI into real work, this distinction becomes practical. A policy can say that sensitive information must not be entered into a public model, that people must remain accountable for consequential decisions, or that suppliers must be assessed. Those are important constraints. But a policy alone cannot answer the questions that arrive later: was this particular use case approved? Was the model’s scope understood? Which data path was accepted? Who agreed the residual risk? What changed after deployment?

    A credible operating model needs a memory.

    Policies set direction. Decision records make it operational.

    Policies are deliberately general. They establish principles, boundaries and responsibilities that should endure beyond one product or project. Decisions are local: a team wants to use a specific model, for a defined purpose, with particular data, controls and owners. Treating the two as the same thing creates a familiar gap. The policy exists, the project moves quickly, and the evidence of interpretation is scattered between a ticket, a meeting, a procurement folder and someone’s memory.

    A lightweight decision record closes that gap without turning governance into a ceremony. It should capture:

    • the decision being made and the business context;
    • the policy, standard or legal obligation that applies;
    • the evidence considered, including supplier material and testing;
    • the accountable owner and any reviewers or approvers;
    • the safeguards, assumptions and residual risks accepted; and
    • a review trigger: a date, material model change, incident, new data source or change in use.

    That is not bureaucracy for its own sake. It makes a decision legible to the next person who has to operate, challenge, audit or improve it. It also stops a generic approval from silently becoming permission for a different system, dataset or purpose six months later.

    Why AI makes the evidence problem sharper

    AI systems shift in ways that ordinary software procurement often does not. A model provider may change behaviour, an integration may begin carrying a new category of data, or users may find a valuable use that was never part of the original assessment. The risk is not just a model producing an odd answer. It is an organisation continuing to rely on an old judgement after the facts supporting that judgement have changed.

    The NIST AI Risk Management Framework frames AI risk management as an ongoing activity across design, development, use and evaluation. Its companion work on generative AI similarly treats risks as contextual rather than as a one-time checklist. That is a useful corrective: governance should be able to show not only that a control was named, but how it was applied to a particular use.

    Decision records are the bridge between that principle and day-to-day work. A supplier questionnaire may show what a vendor said at a point in time. A testing report may show what was observed. A decision record connects those inputs to an accountable judgement: this use was accepted for this purpose, subject to these conditions, until this review event.

    Make review triggers explicit

    The most overlooked field in a decision record is the trigger to reopen it. A date is useful, but it is not enough. Good governance is also event-driven. A decision should return for review when, for example:

    • the model, provider, hosting location or core capability changes;
    • a new data type enters the workflow;
    • the system moves from assistance to automation, or affects a consequential decision;
    • monitoring identifies a meaningful performance, bias, security or privacy concern; or
    • a policy, regulatory expectation or internal risk appetite changes.

    These triggers turn a static assessment into a living control. They also create a more honest conversation with delivery teams: approval is not a permanent green light. It is a bounded judgement, made under stated conditions.

    This is especially valuable in public-sector and other high-consequence settings, where explainability needs to include the organisation’s own choices. Our recent look at UK police use of AI made the point from a different angle: capability can expand faster than the operating model that gives it legitimacy. A durable trail of decisions does not make a difficult use case acceptable by itself, but it makes challenge, oversight and correction possible.

    Keep the record proportionate

    Not every decision needs a board paper. The record should be proportionate to the potential impact. A low-risk internal drafting assistant might require a simple owner, approved data boundary, supplier assessment and annual review. A system that influences eligibility, enforcement, safety, employment or access to services needs much more: clear authority, stronger evidence, testing, affected-party considerations and a route to challenge.

    What matters is consistency. If every team invents its own wording and storage location, leaders cannot see the portfolio of decisions they are carrying. They cannot spot repeated dependencies on one supplier, a control that is repeatedly waived, or a cluster of projects working from the same outdated assumption.

    That is the practical promise of PolicyOps: policies, evidence, decisions, ownership and review should be connected rather than treated as separate administrative tasks. The aim is not to centralise every judgement. It is to make authorised judgement traceable, reviewable and easier to improve.

    Start with one decision that matters

    Teams do not need to redesign their entire governance estate to begin. Pick one live AI use case that has crossed from experiment into recurring work. Identify the accountable owner. Link the relevant policy. Write down the evidence and conditions that support the current decision. Then choose the event that would force the team to revisit it.

    That simple exercise usually reveals the real gaps: an unclear owner, an assumption about data that has never been checked, no agreed route for incidents, or no way to tell whether the supplier has materially changed the service. Those are not paperwork problems. They are operating-model problems.

    AI governance becomes credible when it can answer a modest but demanding question: why are we allowed to do this, and how will we know when that answer is no longer good enough? A well-kept decision record is where that answer lives.

    Further reading

    Continue the AI governance series

  • PolicyOps Starts Where the Policy PDF Ends

    PolicyOps Starts Where the Policy PDF Ends

    Most policy programmes end at the point where the document is published. That is understandable: getting a policy agreed is difficult work. But the PDF is not the operating model. It is the start of one.

    The real test arrives when somebody needs an answer quickly. Can a team use a new AI tool? Which data can go into it? Who can approve an exception? What evidence should be kept? A folder full of policies may contain the answer, but it rarely provides a usable route to it.

    From document to decision

    PolicyOps starts with a simple shift: treat the policy as a source for decisions, not a final destination. That means making the approved version findable, giving people the relevant context, and preserving the path from question to outcome.

    Good policy operations do not turn every question into a compliance theatre exercise. They reduce the friction around routine decisions while making the important ones more visible. If the answer is straightforward, people should be able to see the source and move. If it is uncertain, the escalation path should be clear.

    What a usable policy system needs

    • A source of truth: current, approved material rather than a remembered summary.
    • Context: the relevant clause, related policies and limits behind a short answer.
    • Ownership: a visible route for exceptions and decisions that need judgement.
    • Evidence: enough of a record to explain what happened later.

    AI can make this experience faster, but it should not make authority disappear. A useful assistant can retrieve, summarise and connect material. It cannot quietly become the owner of the decision.

    The practical advantage

    When policy guidance is genuinely operational, people spend less time asking where the rule lives and more time applying it well. That is especially valuable in AI governance, where the questions evolve faster than most policy libraries do.

    PolicyOps is a useful example of the approach: governed answers from approved sources, with evidence and human ownership kept in view. The point is not a cleverer policy PDF. It is a better decision route.

  • AI Governance Needs a Decision Route, Not Just a Policy

    AI Governance Needs a Decision Route, Not Just a Policy

    AI governance is often described as a policy problem. Write the right principles, publish a few guardrails, and the organisation is ready. In practice, that is the easy part.

    The harder question is operational: when somebody asks whether they can use an AI tool, which policy applies, who owns the decision, what evidence supports it, and how will that decision be reviewed later?

    Policies are only useful when they can be used

    Most organisations already have policies covering data protection, security, procurement, acceptable use, risk and records. The problem is that those documents are often spread across folders, intranets and inboxes. Even when the rules are sound, people under time pressure struggle to find the relevant version, understand the context or know who can make the call.

    That gap becomes more visible with AI. New tools arrive quickly, prompts and outputs can cross business boundaries, and a broad policy statement rarely answers a specific operational question. “Use AI responsibly” is a sensible principle; it is not a complete workflow.

    PolicyOps turns guidance into a decision route

    PolicyOps is a useful way to think about that missing layer. Rather than treating a policy library as the end point, it treats policies as living inputs to decisions: locate the approved source, show the relevant context, identify ownership, record review and preserve the evidence behind what happened.

    The outcome is not “the AI decided”. It is a more accountable process in which people can see what information was used and retain responsibility for the judgement. That matters when an answer is challenged, when a policy changes, or when an auditor asks why a particular course of action was taken.

    Good governance does not remove judgement. It makes the route to judgement clearer.

    Four practical tests

    • Can people find the approved source? A useful answer should begin with the current policy, not an untraceable summary.
    • Can they inspect the context? A short answer may be helpful, but people need a route back to the clauses, related documents and limitations behind it.
    • Is ownership visible? Decisions involving risk, exceptions or uncertainty need an identifiable reviewer or escalation path.
    • Can the decision be explained later? Retaining the question, source basis and review trail turns a one-off response into evidence.

    The opportunity is disciplined adoption

    AI can make policy guidance easier to access, summarise and apply. But speed without traceability creates its own risk: stale documents, missing context and confident-looking answers that nobody can properly defend.

    The better ambition is not to automate authority away. It is to make good authority easier to exercise. That means combining useful AI assistance with source-backed guidance, visible human ownership and an evidence trail that holds up after the moment has passed.

    For a practical example of this approach, see PolicyOps, which focuses on governed answers from approved policy sources, evidence and human review.

  • The Spreadsheet Was Never the Problem

    The Spreadsheet Was Never the Problem

    Financial timeline data screen with world map close up
    An employee pastes client data into ChatGPT.

    Maybe it was names. Maybe it was a case summary. Maybe it was a spreadsheet full of “just internal” customer details. Maybe it was done with good intentions: summarise this, clean this up, draft a reply, find the pattern, make my life easier.

    And then someone asks the awkward question:

    Where did that data just go?

    That is the moment AI stops being a shiny productivity tool and becomes a governance problem.

    Not because the employee is stupid. Not because AI is evil. But because most organisations have sleepwalked into AI use without building the boring stuff around it: rules, training, approved tools, audit trails, data classification, and a culture where people know what they can and cannot paste into a chatbot.

    The real issue is not ChatGPT

    The lazy response is to say: “Don’t put client data into ChatGPT.”

    Fine. True. But also useless.

    That is like telling staff “don’t have a data breach” and calling it a cyber strategy.

    The better question is: why did the employee think ChatGPT was the right place for that data in the first place?

    Usually the answer is painfully obvious. Their systems are clunky. Their workload is too high. Their templates are awful. Their managers want more output with fewer people. Then along comes a tool that can summarise, draft, analyse, and tidy up in seconds.

    Of course people use it.

    AI adoption is not always a boardroom strategy. Sometimes it is a knackered employee at 4:45 p.m. trying to make a horrible spreadsheet less horrible.

    “But does OpenAI train on it?”

    That depends on the product and settings.

    OpenAI says that by default it does not train on business data from products such as ChatGPT Business, ChatGPT Enterprise, and the API, unless the organisation opts in. It also says enterprise business data is not used for model training by default.

    But that does not magically solve the problem.

    Because data protection is not only about model training. It is about whether personal or confidential information was shared with an external system lawfully, securely, proportionately, and for a proper purpose. The UK ICO’s AI guidance is clear that organisations using AI with personal data still need to think about UK GDPR principles such as lawfulness, fairness, transparency, accountability, security, and data minimisation.

    So the question is not just:

    “Will the model learn from this?”

    It is also:

    “Should this data have gone there at all?”

    That is the bit people miss.

    The employee is not always the villain

    There is a temptation to turn this into a misconduct story.

    “Employee pasted client data into AI. Employee bad. Problem solved.”

    That may be emotionally satisfying, but it is often organisationally dishonest.

    Because if staff have never been trained, if there is no approved AI tool, if policies are vague, if senior leaders keep saying “we need to innovate,” and if productivity pressure is relentless, then this is not just an individual failure. It is a predictable consequence of unmanaged adoption.

    People do not wait for permission when a tool is useful. They just use it.

    That is exactly why organisations need AI rules before the panic starts.

    What should happen now?

    First, do not pretend it did not happen. Work out what was pasted in, whether it included personal data, commercially sensitive material, legal privilege, client confidential information, special category data, or anything regulated.

    Second, identify what tool was used. A personal free account is a very different risk profile from an approved enterprise environment with contractual controls, admin oversight, data controls, and retention settings.

    Third, assess the harm. Was this a one-off prompt containing low-risk information, or was it a full client dataset? Was the data anonymised? Was it copied from a live system? Could individuals be identified? Was there a duty to notify the client, regulator, or data protection officer?

    Fourth, stop making it about “AI panic” and start making it about data handling. The same employee could have emailed the spreadsheet to the wrong person, uploaded it to a random PDF converter, or put it into an unapproved transcription app. ChatGPT is the headline, but poor data discipline is the disease.

    The policy most organisations actually need

    A good AI policy does not need to be a 90-page legal monument that nobody reads.

    It needs to answer basic questions:

    What AI tools are approved?

    What data can go into them?

    What data is banned?

    When must data be anonymised?

    Who signs off higher-risk use?

    What should staff do if they make a mistake?

    Can AI outputs be used directly, or must a human check them?

    Are staff allowed to use personal AI accounts for work?

    What happens with client, legal, HR, medical, financial, policing, or sensitive operational data?

    The best policy is not “don’t use AI.”

    That will fail.

    The best policy is: use AI here, not there; with this data, not that data; through this approved route, not your personal account at midnight.

    The awkward truth

    Most organisations want AI productivity without AI governance.

    They want the speed, the savings, the slick demos, the innovation strategy, the LinkedIn post about “embracing the future.”

    But they do not want to do the dull work of deciding what data is too sensitive, which systems are approved, who is accountable, and how staff are trained.

    Then, when something goes wrong, they act shocked.

    They should not be shocked.

    This is exactly what happens when powerful tools arrive before proper rules.

    So, now what?

    If an employee pasted client data into ChatGPT, the organisation should investigate it properly, contain any risk, and learn from it.

    But it should also resist the urge to treat the employee as the whole problem.

    The real problem may be that the organisation gave people pressure, tools, ambiguity, and no guardrails — then acted surprised when someone drove into a ditch.

    AI governance is not about killing innovation.

    It is about making sure innovation does not quietly turn into a data breach with better branding.
  • The Met, Palantir, and the Problem With “Trust Us” AI

    The Met, Palantir, and the Problem With “Trust Us” AI

    Police vehicle at night representing AI-assisted policing and public surveillance systems

    There are two lazy ways to talk about AI in policing.

    The first is to say it is obviously sinister. Surveillance. Minority Report. Shadowy tech firms. A giant glowing database deciding who gets nicked next.

    The second is to say it is obviously necessary. Crime is complex. Data is overwhelming. Police are drowning in demand. Technology can help. Stop being dramatic.

    The annoying truth is that both sides have a point.

    That is what makes the row over the Metropolitan Police and Palantir interesting.

    London mayor Sadiq Khan has blocked a proposed £50m deal between the Met and Palantir, the US data analytics company. The Met reportedly wanted to use Palantir’s AI technology to automate intelligence analysis in criminal investigations. City Hall raised concerns about procurement, value for money, legal risk, ethics, reputation, and the fact that Palantir appeared to be the only supplier seriously considered.

    That is not just a procurement story.

    It is the future of public-sector AI arriving in the most British way possible: through a row about process, paperwork, public trust, and whether anyone remembered to make it look legitimate.

    Policing does need better tools

    Let’s start with the uncomfortable bit.

    The Met is not wrong to want better technology.

    Modern policing runs on data. Crime reports, intelligence logs, phone downloads, CCTV, ANPR hits, digital evidence, custody records, case files, officer statements, safeguarding referrals, social media, financial records, body-worn video, call logs, risk assessments, disclosure schedules — endless oceans of information, usually spread across systems that look like they were designed during a hostage negotiation with Microsoft Access.

    Officers are not short of things to read.

    They are short of time, clarity, and usable intelligence.

    That matters. Because buried inside all that admin sludge are patterns, risks, suspects, victims, links, timelines, locations, and warning signs. AI could help find them faster. It could reduce duplication. It could surface connections that humans miss. It could help investigators spend less time wrestling databases and more time actually investigating.

    That is not sinister.

    That is sensible.

    A well-designed AI system in policing could be genuinely useful. Maybe even transformative. Not because it replaces judgement, but because it helps organise chaos.

    And policing has plenty of chaos to organise.

    But “we are busy” is not a governance model

    The problem is that usefulness does not cancel out legitimacy.

    This is where public bodies often go wrong with technology. They start from a real operational problem, find a tool that might help, get excited, and then treat scrutiny as an inconvenience.

    That will not work with AI.

    Especially not in policing.

    The police already hold serious powers over the public. They can stop you, search you, arrest you, seize your property, access your data, build intelligence pictures, and make decisions that can change the direction of your life.

    So when a police force wants to plug powerful AI/data analytics into that system, the answer cannot simply be:

    “Trust us, it’ll be useful.”

    No.

    Show the public the safeguards. Show the audit trail. Show the procurement process. Show the bias testing. Show the human oversight. Show who owns the data. Show who can access it. Show what the system is allowed to do. Show what it is forbidden from doing. Show what happens when it gets things wrong.

    Because it will get things wrong.

    Every system does.

    The question is whether the failure is visible, challengeable, and accountable — or whether it disappears into the comforting fog of “operational sensitivity.”

    Palantir brings baggage

    Palantir is not just any software company.

    It is a major US tech firm with deep links to defence, intelligence, public-sector data projects, and controversial government work. It already has significant UK public contracts, including with the NHS and Ministry of Defence. Critics have also raised concerns about its work connected to immigration enforcement and military uses abroad.

    That does not automatically mean Palantir should be banned from public contracts.

    But it does mean the process has to be cleaner than clean.

    If you are a controversial company selling powerful analytics into policing, you cannot rely on the pitch being “our tech works.” That is not enough. Public trust is not just about technical performance. It is about democratic consent.

    And democratic consent is difficult when the public only hears about the deal once politicians, journalists, campaigners, police leaders, and the company itself are already scrapping over it.

    That is the problem.

    By the time everyone is arguing, the trust has already gone.

    The vendor-lock-in problem

    One of the concerns reportedly raised by City Hall was the risk of becoming locked into Palantir’s technology.

    That may sound dull.

    It is not.

    Vendor lock-in is one of the quiet dangers of public-sector technology. A system is introduced to solve one problem. It becomes embedded. Staff are trained on it. Workflows start depending on it. Data moves into it. Other systems connect to it. Then, a few years later, changing supplier becomes expensive, disruptive, and politically painful.

    At that point, the supplier is not just providing a tool.

    It has become part of the machinery.

    That is a big deal in policing.

    Because once a private company becomes deeply embedded in how intelligence is analysed, how risk is surfaced, how misconduct is spotted, or how investigations are prioritised, it is not merely selling software.

    It is shaping judgement.

    That deserves more than a hurried business case and a few reassuring words about innovation.

    AI aimed at officers is still surveillance

    There is another layer here.

    Reports also say the Met had trialled Palantir technology to identify corrupt or poorly performing officers, with examples including dishonesty around computerised systems and undeclared associations. LBC reported that the Metropolitan Police Federation criticised that use of AI, saying it would damage officer trust and morale, while the Met said the trial identified unacceptable behaviours.

    This is where the debate gets awkward.

    Most people want corrupt police officers found and removed.

    Most police officers want corrupt police officers found and removed.

    But using AI to monitor staff behaviour still raises serious questions. What data is being analysed? What thresholds are being used? Who reviews the outputs? Can officers challenge the conclusions? Are innocent patterns being misread? Is the system detecting misconduct, or just building suspicion from fragments?

    AI does not need to be malicious to be dangerous.

    It only needs to be persuasive.

    A flagged pattern can become a suspicion. A suspicion can become a PSD referral. A referral can become a career-changing process. And even if the AI is wrong, the human brain has already been nudged.

    That is why “human oversight” has to mean more than someone looking at a dashboard and nodding.

    This is not anti-AI. It is anti-bullshit.

    The lazy response is to frame this as pro-police technology versus anti-police politics.

    That misses the point.

    AI in policing could be excellent.

    It could help solve crimes, link offences, reduce admin, spot risk, identify vulnerable victims, manage disclosure, and make investigations less chaotic. Used properly, it could give officers back time and give the public a better service.

    But if the police want AI powers, they need AI legitimacy.

    That means open competition where possible. Clear rules. External scrutiny. Public explanation. Proper data governance. Independent testing. Human accountability. And honesty about what the tool does and does not do.

    Not because everyone is paranoid.

    Because policing depends on consent.

    And consent does not survive secrecy, shortcuts, and “don’t worry, we’ve got this” energy.

    The real lesson

    The Met’s argument is understandable: policing needs to modernise. The data burden is too big, the demand is too high, and old systems are not good enough.

    City Hall’s objection is also understandable: powerful AI in policing cannot be waved through on a questionable procurement process with a controversial supplier and vague reassurances about public safety.

    That tension is not going away.

    In fact, it is probably the future.

    Every public service will face the same problem. Hospitals, councils, schools, courts, benefits systems, immigration, defence, policing. AI will offer real gains. It will also concentrate power, create dependency, and make decisions harder to understand.

    The question is not whether AI belongs in policing.

    It probably does.

    The question is whether policing can introduce AI in a way that earns trust rather than demands it.

    Because “trust us” is not a strategy.

    It is a warning sign.

  • When the AI Policy Hallucinates

    When the AI Policy Hallucinates

    Policy papers and digital analysis concept representing AI governance and document review

    There are some stories that are so neat they almost feel planted.

    South Africa has had to withdraw an early draft of its national AI policy after it was found to contain fictitious and potentially AI-generated references. The policy was meant to help position the country as a leader in artificial intelligence. Instead, it became a case study in one of AI’s most basic risks: it can sound clever while making things up. Reuters reported that an independent panel has now been appointed to review the policy, with a revised version expected for public comment by January 2027.

    That is not just embarrassing.

    It is perfect.

    A government document about regulating AI appears to have been undermined by the exact sort of AI problem the policy should probably have warned about.

    You could not design a cleaner metaphor if you tried.

    The problem is not that AI made a mistake

    AI makes mistakes. That is not news.

    Anyone who has used ChatGPT, Claude, Gemini, Copilot, or any similar tool for more than ten minutes knows this. These systems can be incredibly useful. They can summarise, structure, draft, explain, brainstorm, translate, analyse, and generally act like a very fast assistant with no coffee breaks and no sense of shame.

    But they can also produce complete rubbish with the confidence of a senior consultant billing by the hour.

    That is the real issue.

    Not the error itself.

    The confidence.

    AI does not always say, “I’m not sure.” It often says, “Here you go,” and hands you something that looks finished. It gives you headings, citations, polished language, impressive structure, and the general aroma of competence.

    And because it looks like work, people mistake it for work.

    That is where things go wrong.

    The danger is the handover point

    The South Africa story is not really about whether someone used AI.

    Of course governments will use AI. So will businesses, universities, councils, police forces, law firms, hospitals, journalists, charities, and everyone else currently pretending they are “exploring the technology” while quietly pasting things into chatbots.

    The issue is not use.

    The issue is supervision.

    AI is not dangerous because it drafts. AI is dangerous because people stop checking the draft.

    That is the thin, boring line between productivity and public humiliation.

    If an AI tool creates a reference list, someone still needs to verify the references exist.

    If an AI system summarises evidence, someone still needs to check the evidence.

    If an AI model proposes policy, someone still needs to understand the policy.

    If an AI tool helps write a risk assessment, someone still owns the risk.

    This is the bit that will separate serious organisations from performative ones.

    The serious ones will build verification into the process.

    The performative ones will generate documents faster, publish them sooner, and then act surprised when the wheels come off.

    “Human oversight” cannot just mean a human was nearby

    One of the great phrases of the AI age is going to be human oversight.

    It sounds reassuring. Sensible. Adult.

    But it can mean almost anything.

    A human clicked “approve.”

    A human skimmed the output.

    A human forwarded the document.

    A human sat in the meeting where the thing was discussed.

    That is not oversight. That is scenery.

    Proper human oversight means someone competent has checked the output against reality. It means the human is not just present, but responsible. It means they understand the tool well enough to know where it fails. It means the boring checks still happen.

    Especially the boring checks.

    Because AI failure is often not dramatic.

    It is not always a robot going rogue. Sometimes it is a fake academic paper in a reference list. Sometimes it is a wrong legal citation. Sometimes it is a made-up quote. Sometimes it is a spreadsheet formula that looks fine until it quietly ruins a budget.

    The failures are small until they are not.

    This is why AI literacy matters

    AI literacy does not mean everyone needs to become a machine learning engineer.

    Most people do not need to understand the maths. They do not need to train models, fine-tune transformers, or pretend they know what a vector database is at networking events.

    But they do need to understand the behaviour.

    They need to know that AI can hallucinate.

    They need to know that fluency is not accuracy.

    They need to know that a confident answer is not the same thing as a correct one.

    They need to know that citations, names, dates, case law, policies, academic papers, technical standards, statistics, and quotes are all high-risk areas.

    They need to know when to use AI as a drafting assistant and when to treat it like a suspicious intern who has just discovered Wikipedia and cocaine.

    Helpful? Yes.

    Fast? Definitely.

    Reliable without checking? Absolutely not.

    The irony is funny, but the lesson is serious

    It is easy to laugh at a government AI policy being pulled because of allegedly AI-generated fake references.

    And we should laugh a bit.

    Because come on.

    But the more serious point is that this will not be the last time. In fact, it is probably happening everywhere already. The only difference is whether anyone notices before publication.

    AI is being introduced into systems that already had weak checking, vague accountability, overloaded staff, and a deep institutional love of polished documents nobody properly reads.

    That is fertile ground for artificial confidence.

    The machine produces confident output.

    The organisation performs confident governance.

    The public gets confident language.

    And somewhere underneath it all, nobody has checked whether the source exists.

    The future belongs to people who can verify

    There is a lot of talk about prompt engineering, automation, agents, workflows, and productivity gains.

    Fine. All useful.

    But the underrated skill of the AI age may be verification.

    Can you check the claim?

    Can you trace the source?

    Can you test the output?

    Can you spot when something sounds right but feels thin?

    Can you tell the difference between a useful draft and a dangerous one?

    That is where the value is going to be.

    Not blindly rejecting AI.

    Not blindly trusting it.

    Using it aggressively, but checking it ruthlessly.

    That should be the standard.

    AI is not the problem. Unchecked confidence is.

    This story does not prove that AI should be kept away from government policy. That would be the wrong lesson.

    AI can absolutely help policymakers. It can compare international approaches, summarise consultation responses, identify gaps, model impacts, explain technical concepts, and help turn dense material into something readable.

    That is good.

    But if AI is helping shape the rules of the future, then the people using it need to be better than the tool. They need to bring judgement, scepticism, domain knowledge, and responsibility.

    Otherwise we are not using AI.

    We are laundering guesses through professional formatting.

    The South Africa case is embarrassing, but useful. It gives every organisation a simple warning:

    Before you announce your AI strategy, make sure your AI has not invented the footnotes.

    Because the future may be artificial.

    But the accountability will still be human.