---
title: "FCA Review Shows Why AI Cyber Defence Needs a Remediation Capacity Test"
description: "The FCA’s new review finds that AI-assisted vulnerability discovery is putting pressure on remediation. The practical test is whether a firm can validate, prioritise and safely close findings before it scales the tools."
url: https://artificiallyconfident.com/fca-ai-cyber-defence-remediation-capacity-test/
date: 2026-09-03
modified: 2026-09-03
author: "Andy"
image: https://artificiallyconfident.com/wp-content/uploads/2026/09/fca-ai-cyber-defence-remediation-capacity-test.png
categories: ["AI Risk and Security"]
type: post
lang: en-US
---

# FCA Review Shows Why AI Cyber Defence Needs a Remediation Capacity Test

On 2 September 2026, the Financial Conduct Authority published a [review of frontier AI and cyber resilience](https://www.fca.org.uk/publications/multi-firm-reviews/frontier-ai-cyber-resilience) reporting that firms’ ability to discover vulnerabilities is advancing faster than their capacity to respond. For organisations expanding AI-assisted security work, the immediate question is whether additional findings will produce safer systems or simply a longer queue.

**Evidence note — last checked 3 September 2026, 08:32 BST (Europe/London).** The immediate impact may be limited: this is not an emergency announcement or a new rule. Its operational lesson is material. The discussion below distinguishes reported observations from Artificially Confident’s proposed approach to testing readiness.

## What the publications establish

The FCA describes observations supplied by firms during engagement, not independently measured performance across the whole sector. Firms reported pressure on remediation even after expert reviewers discounted a substantial proportion of model outputs. This is evidence of an operating problem, not a quantified prediction of how many vulnerabilities every organisation will find.

The review expressly introduces no new rules, guidance or regulatory expectations. [Norton Rose Fulbright’s same-day legal commentary](https://www.regulationtomorrow.com/2026/09/fca-multi-firm-frontier-ai-and-cyber-resilience/) also makes that distinction. Its account corroborates the publication and its content; it is not a separate audit of the firms’ results.

The Bank of England published a complementary [discussion of frontier AI harness engineering](https://www.bankofengland.co.uk/research/fintech/artificial-intelligence-consortium/frontier-ai-information-sharing-forum/frontier-ai-harness-engineering) on 2 September. Here, a “harness” means the tools, information, controls and operating environment surrounding a model. The Bank separates access to source code or production artefacts from access to live production systems, and discusses isolated environments and controls around tool use. It too says the publication creates no new supervisory expectations or policy.

These distinctions matter when specifying a pilot. Giving a system a copy of application code for analysis should not silently grant it permission to interact with customer-facing infrastructure. “Access approved” is too imprecise a decision record if it conceals that difference.

## Count the queue, not just the discoveries

Our [earlier analysis of Anthropic’s vulnerability ledger](https://artificiallyconfident.com/anthropic-ai-vulnerability-ledger-security-bottleneck/) examined the security bottleneck around AI-generated findings. The useful next step is to make that bottleneck measurable inside a particular organisation, with its actual staffing, release calendar and dependencies.

Consider a deliberately simplified example, not a figure from either publication. A pilot adds 20 validated findings a working day, while the team can safely close five. With no other arrivals or closures, the unresolved queue grows by 15 daily. A discovery dashboard may look impressive while the oldest unresolved issue becomes progressively older.

Raw counts still need interpretation. Closing five trivial issues does not necessarily outweigh leaving one serious exposure open. Conversely, a growing queue can reflect improved visibility rather than deteriorating security. A useful report needs both the movement of work and a defensible account of what remains exposed.

## Operational consequence: test capacity before expanding scope

Artificially Confident’s recommendation is a bounded capacity exercise before a material expansion of AI-assisted discovery. This is an implementation proposal, not an FCA-prescribed test or assurance certification.

**Set the comparison before running the pilot.** Select one service and record its existing unresolved work, available engineering hours and normal release constraints. Agree what improvement would justify expansion. Without a baseline, additional staffing or an unusually quiet release week could be mistaken for a benefit delivered by the tool.

**Follow the same cases from arrival to outcome.** Give each finding a stable identifier and timestamp each hand-off. Record duplicates against their original case rather than counting them as fresh successes. Keep rejected findings in a separate, reason-coded set. This makes it possible to discover whether the expensive stage is reproduction, assignment, implementation or waiting for a release.

**Introduce a realistic surge without touching live systems.** Use a controlled exercise with a known set of synthetic cases, including incomplete reports and repeated submissions. Ask the receiving team to work through its ordinary process. Measure how quickly it recognises missing evidence and who can resolve an ownership dispute. Do not confuse a tabletop exercise with proof that a production patch works.

**Test the unsuccessful path.** Include a case in which the proposed change fails its checks, the service owner is unavailable, or a dependency cannot be updated immediately. Require an explicit next decision, an accountable owner and a review date. A workflow that demonstrates only uncomplicated closures has not shown how it behaves when the queue becomes difficult.

**Make expansion a separate decision.** At the end of the exercise, compare unresolved-case age, staff time per accepted finding and verified outcomes with the baseline. Identify any work displaced from existing security commitments. The result may justify broader coverage, a narrower search scope or investment in a particular team. It should not automatically justify buying more discovery capacity.

## What a board or service owner should ask next

A concise decision paper can show three views: what the pilot newly revealed, what changed as a result, and what still needs an explicit decision. Keep those views separate. A generated report belongs in the first; a verified change belongs in the second; an unresolved exception belongs in the third.

For organisations dependent on shared technology, our [analysis of frontier AI and financial-stability oversight](https://artificiallyconfident.com/frontier-ai-cyber-risk-financial-stability-oversight/) provides the wider context. A local pilot should make its dependency assumptions visible rather than treating a supplier’s future fix as completed work.

Neither new publication establishes a universal safe throughput or proves that any particular tool reduces incidents. Watch for stronger outcome evidence, and collect it in the pilot: accepted findings, demonstrable risk reduction, time spent and unintended disruption. The useful question is not simply whether AI can find more. It is whether the organisation can explain what happened to the findings it already has.
