Artificially Confident

Artificially Confident

Practical AI, properly examined

Category: Research

Original Artificially Confident research, public-interest evidence reviews and reproducible datasets about AI and its governance.

  • AI in the Investigation Room: The Tools Are Arriving Before the Operating Model Is Complete

    AI in the Investigation Room: The Tools Are Arriving Before the Operating Model Is Complete

    Original research from Artificially Confident: a source-linked audit of how publicly documented UK policing AI projects handle evidence, human judgment, disclosure and operational control.

    Artificially Confident Research · Evidence audit · Checked 19 August 2026

    Research disclosure

    Study design
    Purposive review of eight publicly documented UK policing AI use cases spanning prototypes, trials and operational use.
    What is scored
    The completeness of public evidence—not legality, safety certification, effectiveness or the quality of a force as a whole.
    Method
    Nine criteria scored 0–2, producing 72 evidence judgments with a rationale and official source URL for every score.
    Important limitation
    “Not publicly evidenced” does not establish that a safeguard is absent inside a force or supplier.

    Download the complete evidence workbook (.xlsx)

    The most important question about artificial intelligence in policing is no longer whether the technology will reach the investigation room. It already has.

    AI systems are being developed to summarise case files, analyse digital evidence, identify possible inconsistencies in witness statements, detect patterns in grooming conversations, link concealed online identities, prioritise high-risk suspects and search facial images. In June 2026, the Home Office launched PoliceAI with £75 million over three years to identify, test and scale tools across England and Wales. Its first-year priorities include case-file assistants, disclosure assistants, crime-data integrity, CCTV and digital-media analysis, and image identification. The accompanying government factsheet also promises a public registry of models used operationally and the checks performed before deployment.

    That is an ambitious programme. It is also the beginning of an operating model, not proof that one is complete.

    Artificially Confident reviewed the public evidence for eight UK policing AI use cases. We scored each against nine questions covering operational status, investigative purpose, provenance, human verification, validation, limitations, training, disclosure and operational sovereignty. The study created 72 individual evidence judgments, each tied to an official public source.

    The result is not that UK policing is recklessly deploying useless machines. Several projects demonstrate thoughtful design, honest testing and strong human safeguards. The result is more specific: the tools are becoming visible faster than the complete operating discipline needed to use, scrutinise and sustain them.

    What we studied

    The sample deliberately spans different stages of maturity:

    • a South Yorkshire Police case-file assistant developed to technology readiness level 4;
    • DRAGON-Spotter, which applies linguistic and machine-learning analysis to online-grooming chat logs;
    • an Avon and Somerset Police platform for contact-centre analysis and crime classification;
    • the Traffic Jam risk tool used in an eight-month trafficking trial;
    • PULSE, a prototype for linking concealed online identities;
    • AMULET, a government-owned transparent LLM pipeline for investigative analysis;
    • West Midlands Police models for prioritising potentially high-harm stalking and harassment suspects; and
    • retrospective facial recognition searches using the Police National Database.

    Most project evidence came from the Office of the Police Chief Scientific Adviser’s 2024/25 Police STAR Fund outcomes and 2023/24 outcomes. The facial-recognition assessment used the Home Office’s public guide to police facial recognition.

    This is a purposive sample, not a census of every AI system used by every force. It is a review of public evidence, not an inspection of confidential systems. A score of zero means that a safeguard was not established in the material we reviewed. It does not prove that the safeguard is absent behind the scenes.

    That distinction matters. Transparency can legitimately be constrained in sensitive investigations. But if the public, courts, defence practitioners, regulators and other forces cannot see how a system was tested or governed, external assurance remains limited even when good internal work may exist.

    The headline result: 67% public-evidence coverage

    Across all eight use cases, the sample recorded 97 evidence points from a possible 144: overall coverage of 67%.

    That aggregate conceals a striking split.

    The public material was very good at saying what the tools were for. Investigative role and purpose scored 100%, while operational phase clarity scored 94%. Limitations and error visibility reached 81%, helped by projects that openly described the need for more testing, their false-positive burden or practical staffing constraints.

    The weakest areas were training and user competence at 38%, followed by disclosure, audit and challenge at 31%.

    Those are not secondary administrative details. They are the bridge between a technically interesting system and a defensible investigation.

    An investigator must know what an AI output means, what it does not mean, what source material supports it, how to verify it and when to reject it. A supervisor must be able to reconstruct how it influenced a line of enquiry. Prosecutors and defence teams may need an intelligible record of the system, its configuration, material outputs, discarded outputs and known limitations. If the model changes, the organisation needs to know which cases were affected.

    The Responsible AI Checklist for Policing already points in this direction. It asks whether source data, system workings, evaluations, audits, false positives, false negatives and outputs below a decision threshold will be retained where relevant to disclosure. The checklist is useful precisely because it recognises that evidential use creates obligations well beyond model accuracy.

    The challenge is turning those questions into routine operational controls before national scale-up, not after the first contested case.

    There are good models already

    The strongest-scoring project in our sample was the West Midlands Police high-harm prioritisation work, which received 16 of 18 points.

    Its value was not that the algorithm appeared perfect. It plainly was not. An academic replication found that a high-recall model identified 76% of high-harm suspects while also including 74% of lower-risk people. That is a substantial false-positive burden.

    Publishing the trade-off is a strength. It makes clear that the system is a prioritisation aid, not an oracle. The project also tested the model alongside human judgment in a triage clinic and produced the RUDI framework to support transparency, justification, lawfulness and accountability.

    Retrospective facial recognition scored 14 of 18. The government guide describes a layered process: an algorithm returns candidate images, a specially trained operator assesses them, and an investigating officer conducts additional checks using all the available evidence. The guide also publishes a testing result—99% retrieval of a correct match when one was present at police settings—and acknowledges a limited circumstance in which some demographic groups were more likely to be incorrectly returned.

    Again, the important feature is not the headline percentage alone. It is the combination of testing, a known limitation and a human decision boundary.

    DRAGON-Spotter offers another promising pattern. Its outputs are designed to be cross-referenced with validated forensic tools and checked against wider case data. Domain experts influence how the system is refined. That is what human involvement should mean: not a ceremonial click at the end, but informed control over how an output is produced, interpreted and verified.

    “Human in the loop” is not enough

    The phrase appears throughout responsible-AI discussions because it sounds reassuring. On its own, it proves very little.

    A human can be present but poorly trained. They can be encouraged to trust a confident summary. They may not see the source passages behind a generated conclusion. They may be unable to tell whether an apparent inconsistency is meaningful, whether a model has silently changed or whether an output falls within the tool’s validated conditions.

    For investigative use, a credible human-control protocol should answer at least five questions:

    1. Who is competent and authorised to use the tool?
    2. What source evidence must they examine before acting on an output?
    3. What decisions may never be made from the output alone?
    4. What must be recorded when the investigator accepts, rejects or modifies the output?
    5. Who reviews the decision when the tool affects a significant line of enquiry?

    The College of Policing’s guidance on selecting and working with AI suppliers tells forces to begin with a clear problem, understand their data, assess interoperability and hidden costs, and apply the responsible-AI checklist. That is a good foundation. National adoption now needs a standard competence and supervision layer that can travel with a tool from pilot to deployment.

    The disclosure problem cannot be bolted on later

    In July 2026, the government said AI would be used to help police review and summarise evidence as part of major disclosure reform. It also accepted the case for centralised police technology procurement and a national governance forum for disclosure technology. The announcement reflects a real problem: modern investigations may contain millions of documents and enormous quantities of digital material.

    But disclosure is not merely a volume problem waiting for better search.

    AI can influence what an investigator notices, which material is prioritised, how a document is described and which apparent contradictions are pursued. A generated summary may never become evidence itself while still shaping the investigation profoundly. That makes the tool’s influence part of the investigative history.

    A disclosure-ready system should therefore preserve more than its final answer. Depending on the use case, the record may need to include:

    • the source material presented to the system;
    • the model and version used;
    • relevant configuration, prompts, thresholds and retrieval settings;
    • the output shown to the investigator;
    • links from claims back to source material;
    • warnings, uncertainty and known limitations;
    • the investigator’s acceptance, rejection or modification;
    • material outputs that were dismissed or fell below a threshold;
    • subsequent model changes and any affected-case review; and
    • a form that prosecutors, courts and defence practitioners can understand and challenge.

    That is PolicyOps applied to investigation: governance expressed as part of the work, rather than a policy document stored separately from it.

    UK policing should create space for UK-focused suppliers

    There is a serious case for UK police forces to look more deliberately towards UK-focused companies, university spin-outs and public-interest technical partnerships.

    The case is not that a British postcode makes software trustworthy. It does not. Nor should an international supplier be excluded when it can meet the operational requirement and assurance standard.

    The case is that investigative technology is unusually dependent on jurisdiction, practice and institutional access. A supplier needs to understand the Criminal Procedure and Investigations Act, UK data-protection law, policing data standards, the relationship with the Crown Prosecution Service, force assurance processes, operational security and the practical reality of investigators working across fragmented systems. It must be willing to expose enough technical detail for testing, audit and legal challenge. It should be able to work alongside practitioners over time rather than deliver a generic product and disappear behind a contract.

    Several of the more convincing projects in this review emerged from UK police-university collaboration, government-owned capability or an open-source approach. South Yorkshire’s case-file assistant used licence-free models within police systems and was packaged for reuse. AMULET is government-owned and can work with open-source and locally deployed models. The West Midlands project combined police development with independent replication by UK universities and a freely available implementation framework.

    These are different models, but they share something valuable: operational learning and technical control remain closer to the public institution.

    That proximity can reduce dependence on a single platform, make system changes easier to inspect, improve access to technical experts and support designs built around UK evidential duties from the beginning. It can also strengthen domestic capability in a field where the state will otherwise become a permanent buyer of opaque systems it cannot meaningfully interrogate.

    The National Cyber Security Centre’s machine-learning guidance reinforces the underlying principle: organisations should understand their AI supply chain and require transparency, traceability, validation and verification. Those requirements apply regardless of where a supplier is headquartered.

    A UK policing AI supplier compact

    The practical answer is not a nationality preference masquerading as assurance. It is a published supplier compact that any company can meet, designed so capable UK SMEs are genuinely able to compete.

    The compact should require:

    1. UK-law and workflow fit. The supplier must show how the product supports applicable policing, data-protection, equality, evidential and disclosure duties.
    2. Source-level traceability. Material outputs must link back to the data that supports them, with enough context for an investigator to verify the result.
    3. Independent validation. Performance claims should be tested on representative UK operational conditions, including relevant subgroup and failure-mode analysis.
    4. A defined human decision boundary. Contracts and operating procedures should state what the system may recommend, what a person must check and what the system may never decide alone.
    5. Training and competence. Deployment should include role-based training, assessment, supervision, refresher requirements and records of authorised users.
    6. Disclosure and challenge by design. The system must preserve an intelligible history of material inputs, outputs, settings, decisions and changes.
    7. Secure data control. Forces should know where data is processed, who can access it, whether it is reused for training and how every subcontractor is governed.
    8. Portability and exit. Data, audit logs, prompts, configurations and relevant documentation must be exportable in usable formats. The force must be able to leave without losing its investigative history.
    9. Change and incident control. Model updates, regressions, vulnerabilities and significant errors should trigger recorded review and, where necessary, affected-case analysis.
    10. Proportionate access for smaller suppliers. Suitable procurements should be divided into lots, use realistic timescales and avoid financial or insurance requirements unrelated to the actual risk.

    That final point is not special pleading. Government guidance on the Procurement Act says contracting authorities must consider barriers faced by SMEs and whether they can be reduced. It specifically identifies breaking contracts into lots, publishing pipelines and making engagement more accessible. The official guidance recognises that smaller suppliers can provide innovative public-service solutions but are often excluded by procurement design rather than technical capability.

    Centralised procurement can help policing avoid 43 forces repeatedly solving the same assurance problem. But aggregation must not turn one national requirement into one enormous contract that only a handful of global platforms can bid for. National standards and shared testing should coexist with modular procurement and local innovation.

    The conclusion is not “slow down”

    The evidence does not support a simple story in which police forces do not know how to use AI. It shows a system learning in public, with some genuinely good work and some important gaps.

    The strongest projects are candid about limitations, keep humans close to the evidence and make their stage of maturity clear. The weakest public evidence concerns the disciplines that become most important after a prototype leaves the laboratory: competence, disclosure, challenge, supplier control and long-term operational ownership.

    PoliceAI creates an opportunity to close those gaps nationally before fragmented adoption hardens them into 43 different problems.

    The next phase should treat AI tools as governed investigative capabilities, not clever software licences. It should welcome international technology where it meets the standard, while deliberately cultivating UK-focused companies and research partnerships that can work inside the country’s legal and operational reality.

    The technology is arriving. The job now is to make the operating model arrive with it.

    Method, corrections and reuse

    The workbook contains the scoring rubric, all 72 criterion-level rationales, formulas, source links, coding caveats and reproduction instructions. This page is a dated research record; material corrections will be logged rather than silently substituted.

    Editorial disclosure: Artificially Confident has previously linked to CopPlan, a UK-focused investigation-planning product. CopPlan was not included in or scored by this study because the reviewed official project evidence did not place it within the sampled police AI deployments. No supplier claim was used as scoring evidence.

    For corrections or evidence we may have missed, use the contact page. Our publication standards are explained in the editorial method.

  • UK Police AI Transparency Index: What 10 Forces Disclose

    UK Police AI Transparency Index: What 10 Forces Disclose

    This is the first release in Artificially Confident Research: original, source-linked work designed to make consequential AI systems easier to inspect rather than merely easier to discuss.

    Artificially Confident Research · Pilot study · Evidence checked 19 August 2026

    Research disclosure

    Study design
    Purposive pilot of ten UK territorial police forces with publicly reported AI or algorithmic activity.
    What is scored
    Public disclosure quality for one reference use case per force—not legality, effectiveness or the force as a whole.
    Method
    Eight criteria scored 0–2 using public official sources, with a rationale and URL retained for every score.
    Important limitation
    “No public evidence found” does not mean the underlying governance activity was not performed.

    Download the complete research workbook (.xlsx)

    Police forces are adopting systems that classify messages, search faces, support risk assessment and help staff handle information. The public argument often jumps straight to whether those systems are accurate, biased or lawful.

    There is a more basic question first: can a member of the public find out what a force is using, why it is using it, what role a human plays and how the system is checked?

    Artificially Confident reviewed public information for ten UK territorial police forces. This was a small, purposive pilot rather than a national league table. We selected forces with publicly reported AI or algorithmic activity, chose one reference use case for each, and scored the quality of the disclosure — not the quality of the technology.

    The result is more nuanced than “transparent” or “secretive”. Some forces publish genuinely useful operational records. Others publish responsible-sounding principles without enough tool-specific evidence to let the public test those claims. And the government’s national algorithmic transparency repository does not currently operate as an inventory of police AI.

    What we measured

    We used eight equally weighted questions:

    1. Is the information easy to find?
    2. Does it explain the purpose and operational phase?
    3. Does it explain the human role?
    4. Is ownership or a governance route visible?
    5. Are data and privacy controls explained?
    6. Are risks, limitations, testing or impact assessments disclosed?
    7. Are the supplier, system and technical limits described?
    8. Are monitoring, results, review or challenge routes visible?

    Each question scored zero, one or two. A two required clear, accessible and tool-specific public evidence. A one meant partial, general or fragmented information. A zero meant that we did not find relevant public evidence through the defined official-source search route.

    That last distinction matters. “Not publicly disclosed” is not the same as “not done”. Police forces may have internal assessments that are not published, and operational security can justify withholding some detail. This research is about what the public can verify.

    The full methodology, criterion-level rationales and source URLs are preserved in the research workbook.

    The pilot results

    Force Reference use case Score / 16
    Essex Police Live Facial Recognition 16
    West Yorkshire Police Live Facial Recognition 16
    Metropolitan Police Service Live Facial Recognition 15
    Greater Manchester Police Live Facial Recognition 15
    Hampshire and Isle of Wight Constabulary DARAT 15
    Thames Valley Police DARAT 15
    South Wales Police Operator Initiated Facial Recognition 14
    West Midlands Police Orlo Shield and Assist 13
    Kent Police Live Facial Recognition 12
    Avon and Somerset Police Internal generative AI, including Microsoft Copilot 10

    These numbers should not be read as a ranking of the forces themselves. They are scores for the public disclosure surrounding one selected use case on one date. Facial-recognition deployments have attracted exceptional legal, political and public scrutiny, so it is unsurprising that their documentation is often more developed than disclosure for administrative generative AI.

    That is itself an important result: transparency appears to be driven by the visibility and controversy of a use case, not yet by a consistent force-wide publication system.

    What good disclosure looks like

    The strongest pages behave less like public relations and more like operational records.

    Essex Police’s Live Facial Recognition hub combines planned deployments, a downloadable deployment history, policies, impact assessments and independent research. It explains its operating threshold and discusses different interpretations of 2026 performance studies. That is unusually valuable because it lets a reader see that assurance is not always a single, frictionless answer.

    West Yorkshire Police publishes upcoming and previous deployments, explains the watchlist and deletion process, names the software, describes where a trained operator intervenes and links to policy, impact and legal material.

    Greater Manchester Police similarly explains the human decision point, data deletion, operating contexts and supplier, with linked impact and legal documents.

    The Metropolitan Police facial-recognition hub distinguishes live, retrospective and operator-initiated systems, provides multi-year deployment records and links policy, data-protection, equality and system-performance material.

    These disclosures are not proof that every deployment is correct. They are evidence that members of the public have something concrete to interrogate.

    The oldest national records are still informative — and visibly stale

    The UK government’s Algorithmic Transparency Recording Standard repository contained 64 UK records at the evidence cut-off. Only two police organisations appeared in its organisation filter: Hampshire and Thames Valley Police jointly, and West Midlands Police.

    The joint DARAT record is detailed. It identifies the team, senior responsible owner and developer; describes intended decision pathways; and publishes an extensive set of risks involving bias, fairness, missing data, model drift, feedback loops and system failure.

    But the record describes a pre-deployment system and dates from the early ATRS pilot. The other police record, West Midlands Police’s exploratory analysis of sexual convictions, concerns a one-off analysis that is now retired.

    This means the repository is useful as a disclosure format but not as a current map of police AI. A member of the public cannot use it to answer the simple inventory question: which operational AI systems are police forces using today?

    That is not a breach of the current ATRS mandate. The government’s scope policy makes the standard mandatory for specified central-government bodies, while recommending it across the broader public sector. Police forces are operationally independent and are not currently required to publish ATRS records.

    The fair conclusion is not that forces are non-compliant. It is that the public lacks a consistent, current and central police AI inventory.

    General principles are useful, but they are not evidence of implementation

    Avon and Somerset Police publishes a clear AI principles page. It says AI is subject to governance, impact assessment, legal and ethical review, monitoring and audit. It also states that generative-AI outputs must be checked and that tools such as Microsoft Copilot support internal productivity rather than autonomous operational decisions.

    Those are sensible commitments. The transparency gap is that the page does not provide an inventory, deployment dates, named owners, linked tool-level assessments, test results or monitoring outcomes. Readers are told that controls exist, but are given limited evidence with which to examine how those controls worked for a particular system.

    West Midlands Police provides a stronger tool-specific explanation for Orlo Shield and Assist. It names the supplier, describes message sorting, moderation, drafting, summaries and image indicators, and repeatedly identifies the human review point. The remaining gap is assurance evidence: no tool-specific impact assessment, quantified testing, review date or results report is linked from the disclosure.

    This difference is central to PolicyOps thinking. A policy statement says what should happen. An operational record shows what was decided, by whom, using which evidence, with what limits, and what happened next.

    The missing object is a decision record

    The pilot suggests that police AI transparency does not primarily need more high-level principles. The National Police Chiefs’ Council has already endorsed a Covenant for Using Artificial Intelligence in Policing, placing transparency, fairness and public confidence at the centre of the approach.

    The missing object is a maintained decision record for each material system.

    A useful record would say:

    • what the system is and which operational phase it is in;
    • the decision or workflow it influences;
    • what a human must review and what they can override;
    • who owns the deployment decision;
    • which data sources are used and how long data is retained;
    • what errors, bias and misuse risks were tested;
    • which supplier and model version are in use;
    • what thresholds or meaningful settings apply;
    • what monitoring has found since deployment;
    • when the record was last reviewed; and
    • how a person can ask questions, complain or challenge an outcome.

    Sensitive operational details can be withheld or generalised. The government’s ATRS policy already recognises exemptions and the need to avoid harmful disclosure. But a security exception should be a reasoned field in a record, not a substitute for the record itself.

    This is where a PolicyOps approach becomes practical. The disclosure should be generated from the same governed workflow that approves, reviews and changes the system. Publication then becomes an output of operational governance, rather than an occasional communications exercise assembled after public pressure.

    What should happen next

    This pilot is deliberately small. The next version should expand to all territorial forces, use a pre-registered search protocol, add a second reviewer for a sample of scores and publish a correction log. It should also distinguish three separate measures:

    1. inventory coverage — how many known systems have a public record;
    2. record quality — how complete each disclosure is; and
    3. record freshness — whether the disclosure reflects the current operational system.

    The most important of those may be freshness. A beautifully detailed pre-deployment record can become misleading if it is never updated after the system changes, launches or retires.

    Police use of AI will remain contested. Better transparency will not resolve every disagreement, and it should not be treated as automatic legitimacy. It does something more basic and necessary: it gives the public, oversight bodies and police leaders a shared record of what is actually being operated.

    That is the point at which debate can move from slogans to evidence.

    Research note

    This article reports a purposive ten-force pilot using public official sources checked on 19 August 2026. Scores assess the disclosure for one reference use case per force. They do not assess legality, effectiveness or the totality of a force’s AI use. “No public evidence found” does not mean an activity was not performed. The complete score rationales, source links and methodology are retained in the accompanying research workbook.

    Method, corrections and reuse

    The workbook contains the scoring rubric, all 80 criterion-level rationales, official source links, evidence dates, confidence flags, limitations and reproduction instructions. This page is a dated research record. Material corrections will be logged rather than silently substituted.

    For questions, corrections or evidence we may have missed, use the contact page. Our wider publication standards are explained in the editorial method.