Artificially Confident

Artificially Confident

Practical AI, properly examined

US Agencies Treat Industrial-Scale AI Distillation as a Cross-Platform Security Threat

Written by

in

Abstract model-extraction routes pass through several cloud and API layers towards a smaller reconstructed model while a detection plane correlates the traffic.

Evidence note: Last checked 9 September 2026 at 05:58 BST (Europe/London). The immediate operational impact is concentrated among model providers, cloud platforms and API aggregators. The wider governance lesson is material for any organisation that operates a valuable model or depends on third-party AI access.

The US National Security Agency, Cybersecurity and Infrastructure Security Agency and Federal Bureau of Investigation issued a joint advisory on 8 September warning that six China-based AI companies have conducted industrial-scale campaigns to extract capabilities from US frontier models.

The agencies allege that DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI obtained billions of tokens through millions of exchanges with versions of Claude, GPT, Gemini and Grok from late 2024 onwards. Their announcement and 18-page advisory describe distributed account pools, proxy services, automated request routing and attempts to obtain hidden reasoning traces.

This is an attribution by US security agencies, not a court finding. The public advisory does not disclose the underlying intelligence in enough detail for outsiders to reproduce every attribution. Reuters reported that China’s embassy in Washington did not immediately respond to a request for comment. Previous Chinese government responses have rejected similar US accusations as groundless. The six companies’ individual responses to this advisory were not established in the sources reviewed.

The advisory describes an access and telemetry problem

Knowledge distillation is not inherently malicious. A developer can train a smaller model on outputs from a larger model to reduce cost or create a specialist system. Frontier laboratories use the technique within their own model families. The agencies’ concern is unauthorised extraction at scale, using access methods designed to evade contractual, geographic and technical controls.

The advisory describes four operating mechanisms. Account pools are used to spread requests and avoid quotas. Central routing systems switch traffic between native APIs, cloud services, aggregators and third-party relays. Metadata is sanitised to remove organisational identifiers. Automated quality checks are then used to detect whether a provider has degraded or altered its responses.

That combination matters because a provider may see only a small part of the campaign. One account, IP address or API route can look unusual without proving coordinated extraction. Correlation across identity, payment, timing, prompt similarity, throughput and model-switching behaviour is needed to reveal the larger operation.

The new warning is supported by provider reporting. In February, Anthropic said it had identified more than 16 million exchanges through about 24,000 fraudulent accounts associated with campaigns it attributed to DeepSeek, Moonshot and MiniMax. A separate Google Threat Intelligence report published on 8 September says some recent extraction campaigns against its models exceeded 100 million prompts and rotated requests across thousands of compromised credentials and fraudulent accounts.

Single-provider controls are not enough

The agencies recommend behavioural monitoring, stronger identity verification, rate limits, targeted response changes and cross-organisation intelligence sharing. The cross-platform element is the most consequential. If campaigns can fail over between providers and hide behind aggregators, each organisation’s view is incomplete by design.

Model security therefore needs evidence that can be compared without exposing customer prompts indiscriminately. Providers and intermediaries need agreed indicators, retention rules, confidence thresholds and channels for sharing infrastructure and behavioural signals. They also need a clear boundary between threat intelligence and competitive information. A useful security programme cannot become a pretext for exchanging sensitive commercial data or blocking legitimate research.

This extends the assurance argument in our review of NIST’s AI cybersecurity guidance: an inventory and risk assessment are only a starting point. Organisations must connect model access, identity, telemetry, incident handling and supplier relationships into one operating control.

Quietly degrading responses creates a governance obligation

The advisory’s most sensitive recommendation is to alter responses for high-confidence distillation activity without telling the suspected operator. Suggested techniques include routing requests to a less capable model, reducing reasoning depth or introducing variations that lower the value of extracted training data.

There is a defensible security rationale. Warning a confirmed adversary would help it improve evasion and identify which data to discard. The same control can also affect legitimate customers, independent evaluators and safety researchers if the detection threshold is wrong.

Any provider using selective degradation should therefore record the evidence, confidence score, approving authority, affected services, review date and appeal route. Researchers and third-party evaluators need disclosure when model behaviour has been changed, as the advisory itself recognises. Otherwise, benchmark results and safety assessments may describe a hidden fallback system rather than the service being evaluated.

The lesson resembles the one from the OpenAI and Hugging Face incident: an AI security control is not only a classifier or technical boundary. It is a decision process with evidence, authority, escalation and external consequences.

What model operators should do next

Organisations that expose valuable models through applications or APIs should test whether they can detect coordinated extraction across accounts and routes, rather than relying on per-user rate limits. They should ask aggregators and cloud partners what telemetry is retained, how suspicious campaigns are correlated and how false positives are reviewed.

Teams procuring custom or fine-tuned models should also add provenance questions to supplier assurance. A vendor should be able to explain its training inputs, permitted use of teacher-model outputs, account controls and response to extraction attempts. Claims about independent capability need supporting evidence when model behaviour may have been acquired through another provider’s service.

What follows will matter as much as the advisory itself: whether the agencies release indicators that other providers can use, whether the named companies issue evidence-backed responses, and whether proposed information-sharing arrangements include safeguards for legitimate users. The immediate task is technical detection. The durable task is to make attribution and countermeasures reviewable before they affect access, research or competition.

How we work: articles are source-led, AI-assisted and editorially reviewed. Read our editorial method.

Reader response

Questions, corrections or a story lead?

Send us a message with enough context to make it useful. Your note will reach the Artificially Confident editorial inbox.