Original research from Artificially Confident: a source-linked audit of how publicly documented UK policing AI projects handle evidence, human judgment, disclosure and operational control.
Artificially Confident Research · Evidence audit · Checked 19 August 2026
Research disclosure
- Study design
- Purposive review of eight publicly documented UK policing AI use cases spanning prototypes, trials and operational use.
- What is scored
- The completeness of public evidence—not legality, safety certification, effectiveness or the quality of a force as a whole.
- Method
- Nine criteria scored 0–2, producing 72 evidence judgments with a rationale and official source URL for every score.
- Important limitation
- “Not publicly evidenced” does not establish that a safeguard is absent inside a force or supplier.
The most important question about artificial intelligence in policing is no longer whether the technology will reach the investigation room. It already has.
AI systems are being developed to summarise case files, analyse digital evidence, identify possible inconsistencies in witness statements, detect patterns in grooming conversations, link concealed online identities, prioritise high-risk suspects and search facial images. In June 2026, the Home Office launched PoliceAI with £75 million over three years to identify, test and scale tools across England and Wales. Its first-year priorities include case-file assistants, disclosure assistants, crime-data integrity, CCTV and digital-media analysis, and image identification. The accompanying government factsheet also promises a public registry of models used operationally and the checks performed before deployment.
That is an ambitious programme. It is also the beginning of an operating model, not proof that one is complete.
Artificially Confident reviewed the public evidence for eight UK policing AI use cases. We scored each against nine questions covering operational status, investigative purpose, provenance, human verification, validation, limitations, training, disclosure and operational sovereignty. The study created 72 individual evidence judgments, each tied to an official public source.
The result is not that UK policing is recklessly deploying useless machines. Several projects demonstrate thoughtful design, honest testing and strong human safeguards. The result is more specific: the tools are becoming visible faster than the complete operating discipline needed to use, scrutinise and sustain them.
What we studied
The sample deliberately spans different stages of maturity:
- a South Yorkshire Police case-file assistant developed to technology readiness level 4;
- DRAGON-Spotter, which applies linguistic and machine-learning analysis to online-grooming chat logs;
- an Avon and Somerset Police platform for contact-centre analysis and crime classification;
- the Traffic Jam risk tool used in an eight-month trafficking trial;
- PULSE, a prototype for linking concealed online identities;
- AMULET, a government-owned transparent LLM pipeline for investigative analysis;
- West Midlands Police models for prioritising potentially high-harm stalking and harassment suspects; and
- retrospective facial recognition searches using the Police National Database.
Most project evidence came from the Office of the Police Chief Scientific Adviser’s 2024/25 Police STAR Fund outcomes and 2023/24 outcomes. The facial-recognition assessment used the Home Office’s public guide to police facial recognition.
This is a purposive sample, not a census of every AI system used by every force. It is a review of public evidence, not an inspection of confidential systems. A score of zero means that a safeguard was not established in the material we reviewed. It does not prove that the safeguard is absent behind the scenes.
That distinction matters. Transparency can legitimately be constrained in sensitive investigations. But if the public, courts, defence practitioners, regulators and other forces cannot see how a system was tested or governed, external assurance remains limited even when good internal work may exist.
The headline result: 67% public-evidence coverage
Across all eight use cases, the sample recorded 97 evidence points from a possible 144: overall coverage of 67%.
That aggregate conceals a striking split.
The public material was very good at saying what the tools were for. Investigative role and purpose scored 100%, while operational phase clarity scored 94%. Limitations and error visibility reached 81%, helped by projects that openly described the need for more testing, their false-positive burden or practical staffing constraints.
The weakest areas were training and user competence at 38%, followed by disclosure, audit and challenge at 31%.
Those are not secondary administrative details. They are the bridge between a technically interesting system and a defensible investigation.
An investigator must know what an AI output means, what it does not mean, what source material supports it, how to verify it and when to reject it. A supervisor must be able to reconstruct how it influenced a line of enquiry. Prosecutors and defence teams may need an intelligible record of the system, its configuration, material outputs, discarded outputs and known limitations. If the model changes, the organisation needs to know which cases were affected.
The Responsible AI Checklist for Policing already points in this direction. It asks whether source data, system workings, evaluations, audits, false positives, false negatives and outputs below a decision threshold will be retained where relevant to disclosure. The checklist is useful precisely because it recognises that evidential use creates obligations well beyond model accuracy.
The challenge is turning those questions into routine operational controls before national scale-up, not after the first contested case.
There are good models already
The strongest-scoring project in our sample was the West Midlands Police high-harm prioritisation work, which received 16 of 18 points.
Its value was not that the algorithm appeared perfect. It plainly was not. An academic replication found that a high-recall model identified 76% of high-harm suspects while also including 74% of lower-risk people. That is a substantial false-positive burden.
Publishing the trade-off is a strength. It makes clear that the system is a prioritisation aid, not an oracle. The project also tested the model alongside human judgment in a triage clinic and produced the RUDI framework to support transparency, justification, lawfulness and accountability.
Retrospective facial recognition scored 14 of 18. The government guide describes a layered process: an algorithm returns candidate images, a specially trained operator assesses them, and an investigating officer conducts additional checks using all the available evidence. The guide also publishes a testing result—99% retrieval of a correct match when one was present at police settings—and acknowledges a limited circumstance in which some demographic groups were more likely to be incorrectly returned.
Again, the important feature is not the headline percentage alone. It is the combination of testing, a known limitation and a human decision boundary.
DRAGON-Spotter offers another promising pattern. Its outputs are designed to be cross-referenced with validated forensic tools and checked against wider case data. Domain experts influence how the system is refined. That is what human involvement should mean: not a ceremonial click at the end, but informed control over how an output is produced, interpreted and verified.
“Human in the loop” is not enough
The phrase appears throughout responsible-AI discussions because it sounds reassuring. On its own, it proves very little.
A human can be present but poorly trained. They can be encouraged to trust a confident summary. They may not see the source passages behind a generated conclusion. They may be unable to tell whether an apparent inconsistency is meaningful, whether a model has silently changed or whether an output falls within the tool’s validated conditions.
For investigative use, a credible human-control protocol should answer at least five questions:
- Who is competent and authorised to use the tool?
- What source evidence must they examine before acting on an output?
- What decisions may never be made from the output alone?
- What must be recorded when the investigator accepts, rejects or modifies the output?
- Who reviews the decision when the tool affects a significant line of enquiry?
The College of Policing’s guidance on selecting and working with AI suppliers tells forces to begin with a clear problem, understand their data, assess interoperability and hidden costs, and apply the responsible-AI checklist. That is a good foundation. National adoption now needs a standard competence and supervision layer that can travel with a tool from pilot to deployment.
The disclosure problem cannot be bolted on later
In July 2026, the government said AI would be used to help police review and summarise evidence as part of major disclosure reform. It also accepted the case for centralised police technology procurement and a national governance forum for disclosure technology. The announcement reflects a real problem: modern investigations may contain millions of documents and enormous quantities of digital material.
But disclosure is not merely a volume problem waiting for better search.
AI can influence what an investigator notices, which material is prioritised, how a document is described and which apparent contradictions are pursued. A generated summary may never become evidence itself while still shaping the investigation profoundly. That makes the tool’s influence part of the investigative history.
A disclosure-ready system should therefore preserve more than its final answer. Depending on the use case, the record may need to include:
- the source material presented to the system;
- the model and version used;
- relevant configuration, prompts, thresholds and retrieval settings;
- the output shown to the investigator;
- links from claims back to source material;
- warnings, uncertainty and known limitations;
- the investigator’s acceptance, rejection or modification;
- material outputs that were dismissed or fell below a threshold;
- subsequent model changes and any affected-case review; and
- a form that prosecutors, courts and defence practitioners can understand and challenge.
That is PolicyOps applied to investigation: governance expressed as part of the work, rather than a policy document stored separately from it.
UK policing should create space for UK-focused suppliers
There is a serious case for UK police forces to look more deliberately towards UK-focused companies, university spin-outs and public-interest technical partnerships.
The case is not that a British postcode makes software trustworthy. It does not. Nor should an international supplier be excluded when it can meet the operational requirement and assurance standard.
The case is that investigative technology is unusually dependent on jurisdiction, practice and institutional access. A supplier needs to understand the Criminal Procedure and Investigations Act, UK data-protection law, policing data standards, the relationship with the Crown Prosecution Service, force assurance processes, operational security and the practical reality of investigators working across fragmented systems. It must be willing to expose enough technical detail for testing, audit and legal challenge. It should be able to work alongside practitioners over time rather than deliver a generic product and disappear behind a contract.
Several of the more convincing projects in this review emerged from UK police-university collaboration, government-owned capability or an open-source approach. South Yorkshire’s case-file assistant used licence-free models within police systems and was packaged for reuse. AMULET is government-owned and can work with open-source and locally deployed models. The West Midlands project combined police development with independent replication by UK universities and a freely available implementation framework.
These are different models, but they share something valuable: operational learning and technical control remain closer to the public institution.
That proximity can reduce dependence on a single platform, make system changes easier to inspect, improve access to technical experts and support designs built around UK evidential duties from the beginning. It can also strengthen domestic capability in a field where the state will otherwise become a permanent buyer of opaque systems it cannot meaningfully interrogate.
The National Cyber Security Centre’s machine-learning guidance reinforces the underlying principle: organisations should understand their AI supply chain and require transparency, traceability, validation and verification. Those requirements apply regardless of where a supplier is headquartered.
A UK policing AI supplier compact
The practical answer is not a nationality preference masquerading as assurance. It is a published supplier compact that any company can meet, designed so capable UK SMEs are genuinely able to compete.
The compact should require:
- UK-law and workflow fit. The supplier must show how the product supports applicable policing, data-protection, equality, evidential and disclosure duties.
- Source-level traceability. Material outputs must link back to the data that supports them, with enough context for an investigator to verify the result.
- Independent validation. Performance claims should be tested on representative UK operational conditions, including relevant subgroup and failure-mode analysis.
- A defined human decision boundary. Contracts and operating procedures should state what the system may recommend, what a person must check and what the system may never decide alone.
- Training and competence. Deployment should include role-based training, assessment, supervision, refresher requirements and records of authorised users.
- Disclosure and challenge by design. The system must preserve an intelligible history of material inputs, outputs, settings, decisions and changes.
- Secure data control. Forces should know where data is processed, who can access it, whether it is reused for training and how every subcontractor is governed.
- Portability and exit. Data, audit logs, prompts, configurations and relevant documentation must be exportable in usable formats. The force must be able to leave without losing its investigative history.
- Change and incident control. Model updates, regressions, vulnerabilities and significant errors should trigger recorded review and, where necessary, affected-case analysis.
- Proportionate access for smaller suppliers. Suitable procurements should be divided into lots, use realistic timescales and avoid financial or insurance requirements unrelated to the actual risk.
That final point is not special pleading. Government guidance on the Procurement Act says contracting authorities must consider barriers faced by SMEs and whether they can be reduced. It specifically identifies breaking contracts into lots, publishing pipelines and making engagement more accessible. The official guidance recognises that smaller suppliers can provide innovative public-service solutions but are often excluded by procurement design rather than technical capability.
Centralised procurement can help policing avoid 43 forces repeatedly solving the same assurance problem. But aggregation must not turn one national requirement into one enormous contract that only a handful of global platforms can bid for. National standards and shared testing should coexist with modular procurement and local innovation.
The conclusion is not “slow down”
The evidence does not support a simple story in which police forces do not know how to use AI. It shows a system learning in public, with some genuinely good work and some important gaps.
The strongest projects are candid about limitations, keep humans close to the evidence and make their stage of maturity clear. The weakest public evidence concerns the disciplines that become most important after a prototype leaves the laboratory: competence, disclosure, challenge, supplier control and long-term operational ownership.
PoliceAI creates an opportunity to close those gaps nationally before fragmented adoption hardens them into 43 different problems.
The next phase should treat AI tools as governed investigative capabilities, not clever software licences. It should welcome international technology where it meets the standard, while deliberately cultivating UK-focused companies and research partnerships that can work inside the country’s legal and operational reality.
The technology is arriving. The job now is to make the operating model arrive with it.
Method, corrections and reuse
The workbook contains the scoring rubric, all 72 criterion-level rationales, formulas, source links, coding caveats and reproduction instructions. This page is a dated research record; material corrections will be logged rather than silently substituted.
Editorial disclosure: Artificially Confident has previously linked to CopPlan, a UK-focused investigation-planning product. CopPlan was not included in or scored by this study because the reviewed official project evidence did not place it within the sampled police AI deployments. No supplier claim was used as scoring evidence.
For corrections or evidence we may have missed, use the contact page. Our publication standards are explained in the editorial method.


