Opinion and futures analysis. This article draws an argument from published evidence about AI-assisted research, mathematics, cybersecurity and governance. The forecasts below are identified as forecasts. They do not assume that every organisation will adopt AI at the same rate or face the same risks.
AI systems are starting to produce research code, mathematical proofs, security findings and policy analysis faster than institutions can reproduce, challenge, attribute or approve them. The next constraint on useful AI adoption will not be access to intelligence. It will be the capacity to verify what that intelligence produces.
OpenAI reported on 6 September 2026 that its research organisation was using 3.1 agent-workdays for every human workday by mid-August. Two days later, it published a claimed solution to the Navier-Stokes Millennium Prize problem with a Lean formalisation. The Financial Conduct Authority has meanwhile reported that frontier models can identify vulnerabilities faster than firms can validate and repair them. These are different domains, but they expose the same operating problem.
Generation scales through compute. Acceptance scales through evidence, qualified attention and authority. Those resources do not expand at the same speed.
Generated output and accepted work
OpenAI’s research-acceleration disclosure gives a useful view of the change. The company says its researchers are writing code faster and running more experiments as agent use rises. It also states that people still set priorities, judge results and decide whether to scale, pause or deploy systems. For successful tasks estimated to take a person four to eight hours, more than half involved at least one human intervention during the six months covered by the analysis.
That distinction separates output from accepted work. Code is not integrated because it exists. An experiment is not knowledge because it ran. A policy summary is not a decision because it reads well.
Our analysis of OpenAI’s research-agent figures called for an acceptance funnel: activity, output retained after review, and the human and control cost of reaching that result. Without those measures, rising agent activity can hide a queue of untested claims and unfinished decisions.
The Navier-Stokes announcement makes the point in a harder setting. OpenAI says an internal system produced an analytical proof and a Lean formalisation. The Clay Mathematics Institute still lists the problem as unsolved. That is not a contradiction or a verdict on the proof. It reflects the difference between submitting evidence and earning acceptance through expert scrutiny, publication and the Institute’s rules.
As our developing analysis of the claimed proof argued, formalisation can test whether a chain of reasoning follows from its stated assumptions. It does not establish that the formal statement matches the intended problem, that the imported definitions are appropriate, or that credit and provenance have been settled.
Verification debt
Organisations already understand technical debt: choices that make delivery faster now but create later repair work. AI introduces a related liability. Verification debt is the review work created when generated output enters an organisation faster than qualified people and controls can check it.
The debt appears in rejected security findings, code waiting for tests, citations nobody has opened, model-written procedures that do not match the governing policy, and decisions whose evidence cannot be reconstructed. Each item may look small. The queue becomes material when staff start treating age, confidence scores or polished language as substitutes for verification.
Cybersecurity offers an early warning. In its 2 September 2026 review, the FCA said firms were finding vulnerabilities faster than their remediation processes could respond. Some models produced large numbers of technically possible findings that were hard to validate, prioritise or act upon. The regulator’s list of constraints included validation capacity, engineering resources, patch testing, change controls and evidence of closure.
That is verification debt with an operational consequence. A larger register of plausible vulnerabilities does not make a bank safer if engineers cannot establish exploitability, patch the affected service and prove that the fix worked. Our review of the FCA findings treated remediation capacity as the real test of AI-enabled defence.
Limits of automated verification
The strongest objection is that capable AI will also automate verification. It will. Agents can run tests, compare documents, reproduce calculations, trace citations and search for contradictory evidence. That should reduce some review costs and catch errors people miss.
Automation does not remove the need for authority. A model can test whether software matches a specification, but it cannot grant the specification legitimacy. It can compare a procedure with a policy library, but the organisation still has to establish which policy governs. It can score competing explanations, but a court, regulator, journal, safety committee or accountable executive must decide what standard of proof applies.
NIST’s draft guide to using AI for Cybersecurity Framework analysis illustrates this boundary. Its worked examples use AI to draft governance reviews and current-state profiles from organisational evidence. NIST describes the examples as possible approaches, not assessment or assurance methods. The generated artefact begins the review; it does not complete it.
AI-on-AI checking also creates correlated failure risk. A generator and reviewer trained on similar data, using the same tools and working from the same mistaken premise can agree with each other. Agreement is useful evidence only when the checking process introduces meaningful independence through different methods, sources, incentives or accountable reviewers.
Forecasts for verification capacity
Verification budgets will become visible. Organisations will track review hours, unresolved findings, reproduction time and evidence gaps alongside model use and token spend. A team that generates more output without expanding its acceptance capacity has increased its queue, not its throughput.
Provenance will become part of the product. Buyers will ask for the model and tool versions, source records, intermediate actions, human corrections and acceptance criteria behind an output. Products that preserve this chain will be easier to approve in regulated and safety-sensitive work than products that deliver only a final answer.
Independent assurance will become a market. Organisations will pay specialists to reproduce AI-generated research, challenge automated risk assessments and test whether internal review controls work. The valuable service will not be a generic AI audit badge. It will be a scoped conclusion tied to named evidence, a defined method and a reviewer who can defend the result.
Review queues will become a risk measure. Boards already see unresolved audit actions and overdue vulnerabilities. AI-generated findings, decisions and code changes will join that reporting. Queue age, severity, reviewer competence and time to closure will matter more than the raw volume produced.
Correction speed will matter more than prediction theatre. No organisation will anticipate every model failure or false claim. A defensible operating model will preserve the original evidence, identify who relied on it, stop further use, publish a dated correction and show what changed. This is less attractive than a claim of perfect foresight, but it is testable.
Management questions for AI-generated work
Leaders do not need a new committee for every model. They need a view of the work waiting to be trusted. The following questions expose the constraint:
- Which AI-generated artefacts are accumulating faster than they are reviewed?
- Who is qualified and authorised to accept each type of artefact?
- What source, model, tool and human-intervention records survive after delivery?
- How long does validation take, and which parts can be automated without removing independence?
- Which signal can stop reliance on a result before it causes a downstream decision?
- Can another reviewer reconstruct why an output was accepted?
- What happens when the review queue exceeds available capacity?
The organisations that answer those questions will not use less AI. They will know which outputs have earned reliance and which remain proposals. As generation becomes abundant, that distinction will decide whether AI increases institutional capability or fills institutions with work they cannot responsibly accept.

