Sources#
- Detecting and countering misuse of AI: September 2026
- GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AI
Summary#
Between January and July 2026, Anthropic disrupted a set of operations in which state-aligned actors, state-linked contractors and commercial spyware vendors used Claude to build and run surveillance systems — actors from China, Iran and West Africa, plus the commercial surveillance-for-hire market, "rang[ing] from operations carried out by a single individual to entire teams." In nearly every case the targets were "the same diaspora and dissident communities these regimes have historically targeted": pro-democracy figures in Hong Kong, Tibetan and Falun Gong communities, Uyghur diaspora and media, Iranian minorities and opposition abroad.
case-study, first-party, no external verification. The disruption claims in particular are the vendor grading its own enforcement — and this is the one section of the report where that grading is visibly, usefully negative.
Three substitutions, which are the actual finding#
Anthropic states three trends, and they are three different institutional functions the model replaced:
1. The engineering workforce. A single Claude subscriber — "likely a Bamako-based independent consultant working with Mali's state intelligence service" — used Claude as "the primary engineering workforce" for Lakana 360, a population-scale domestic interception platform covering "roughly 25 million SIM cards on all three of the country's national mobile operators." Also in this class: an Iranian provincial unit that used Claude as "its engineering department," shipping a malicious Firefox extension (al-Najm al-thāqib, disguised as a prayer-times utility) to production to mass-harvest identities from social networks.
2. The analyst desk. "A religious affairs intelligence collection unit in the PRC that once comprised many teams of analysts has been reduced to a single office, using an AI assistant to produce thousands of investigations per month" (GTG-14020). The operator ran four concurrent workstreams — Catholic cardinals across Asia, the leadership of the Presbyterian Church in Taiwan, Tibetan civil society and the administration in exile, Falun Gong and its affiliated media — against internal templates that required, for each subject, their China-related activities, scandals, and 「抓手」 (zhuāshǒu), the United Front Work Department's term for exploitable leverage. Collection went down to birth dates, birthplaces, immigration dates, social handles, and reconnaissance of religious venues "including floor plans, facades, and structural diagrams."
3. The bureaucracy itself. "One PRC state security bureau used Claude to produce an internal manual on how to use AI in surveillance operations" (GTG-14021) — including the prompt formula for instructing Claude to role-play an intelligence analyst serving the state, written up for distribution to the rest of the bureau. And GTG-14022 ran a version-controlled "public opinion monitoring" framework at v2.6, with a master control table and appendices, processing 15 to 30+ foreign news articles a day into government-format 舆情简报 briefings through Claude's code-execution environment. Anthropic's own verdict on that one is the sentence to keep: "Rather than displaying any novel capabilities, this activity was unique in how it was used as part of the bureaucratic apparatus."
The signature that distinguishes these from opportunistic misuse is therefore not capability but cadence and format: a daily reporting cycle, government document templates, a versioned framework, an institutional how-to manual. Two Iranian units "that shared no code or personnel independently used Claude to solve the same technical and usability issues with a state-run centralized surveillance case management system," which is what integration into a bureaucracy looks like from the outside — the same friction hit twice by people who never met.
A Russian instance, from a second provider (2026-09-24). GTIG's Q2 2026 tracker (GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AI, case-study, first-party) adds the country missing from Anthropic's list. UNC5792, a Russia-based group, put AI models inside automated Telegram monitoring bots that sort channel messages as suspicious or neutral for "information of interest to Russian authorities, such as security threats and extremist content," and write the output as structured intelligence reports. GTIG notes the group already ran a Telegram bot. The AI was added to an existing collection loop to handle classification and reporting, which is the analyst-desk substitution (2 above) in a different state, observed from a different vendor. The same report has PRC-nexus RAVINE CASTLE (APT24) turning exfiltrated data into structured intelligence reports by translation and summarization, and CALANQUE ION (APT42) summarizing exfiltrated data. So "the model writes the intelligence product" is now reported by two vendors across three states.
The cases#
| Case | Actor | What Claude did |
|---|---|---|
| GTG-54009 | "S2T Unlocking Cyberspace," an Israeli-Singaporean commercial intelligence vendor (per open-source research) | Classified batches of ~25 social posts at a time into demographic group, location and political leaning with confidence ratings; sorted diaspora users as pro- or anti-government across six coded groups; produced formal-Arabic briefings styled as government communications with sentiment scores by Gulf nationality and recommended counternarratives. A secondary workstream built 255+ synthetic social accounts as a stock for later. Caught at pilot stage; independently corroborates a 2023 Forbidden Stories investigation of an S2T brochure found in leaked Colombian military files |
| GTG-14010 | PRC-aligned contractor (low confidence), Chinese-language | Converted chatter bulk-extracted from 100+ monitored WhatsApp groups and dozens of Telegram channels into structured Chinese-language target data, profiling individuals by financial stress, family separation and ideological disillusionment — and specifically identifying targets with family remaining in Xinjiang, leverage actionable only through PRC domestic security. Then ran a multi-day covert recruitment of Uyghurs in Syria in Syrian Arabic dialect by an operator with no Arabic, with Claude translating replies in real time and role-playing an "expert" consultant to quality-check the deception. Declined: covert interrogation, large-scale persona cultivation |
| GTG-14020 | PRC religious-affairs collection desk; one user disclosed being a state information security officer | The analyst-desk substitution above; daily reporting across domestic and foreign platforms; prompts instructing Claude to adopt "China's standpoint," call the Tibetan administration in exile an "illegal separatist administration," and apply the state's "evil cult" designation to Falun Gong |
| GTG-14021 | Municipal public and state security organs: a cyber police officer, a police academy student, a bureau | Claude Code with custom skills driving a sentiment-monitoring pipeline that queried a government surveillance database and distributed daily reports to supervisors; templated 'situational awareness' briefings; the institutional AI manual; and pre-operational venue intelligence on lawful overseas protests — the gathering point, route and terminus of a Vancouver pro-democracy march, Uyghur event venues in Turkey, Oslo Freedom Forum screenings |
| GTG-14022 | Commercial contractor for PRC government clients (medium confidence) | The versioned 舆情 briefing pipeline above; political-sensitivity scoring; enforced terminology ('Taiwan government' → 'Taiwan authorities', scare quotes around 'human rights violations'); documents recommending enforcement actions only a state could take, framed in 'three warfares' doctrine |
| GTG-34007 | Two linked Iranian paramilitary/domestic-security units, 16 accounts | A seven-department organization claiming an identity-record database of Iranian nationals and 6,388 Iranians profiled in a single year; social-network analysis over 155,216 tweets naming 39 opposition and diaspora accounts; a messenger de-anonymizer, phone-number-to-identity resolver, national-ID phishing page, Telegram mass-report bot, and the shipped Firefox identity harvester; a web front end for 'Arman', the shared case-management system whose per-subject file holds national ID, beliefs, criminal record, social accounts and an 'action' tab |
| GTG-30004 / 30005 / 30006 | Iran-nexus | An automated OSINT identity-profiling harness over hundreds of Israeli and Jewish-diaspora individuals; open-source naval targeting handbooks against US forces with a scraped personnel roster and CVE research on maritime VSAT and shipboard systems; a domestic mass-surveillance design combining ALPR with mobile-identifier interception; and the SECOMS64 implant family with a Telegram-focused keylogger |
| GTG-50027 | One consultant, Mali's Agence Nationale de la Sécurité d'État | Lakana 360: nationwide call-record, SMS and voice capture across ~25M SIMs; cross-SIM voiceprint tracking that defeats burner-SIM self-protection; flagging of encryption and VPN users; inference of clandestine meetings; geofenced watchlists; joins against the national biometric civil registry. "The warrant requirement was removed, at the operator's request," from the component that writes an LLM-generated dossier on any phone number, which was reclassified as a national pipeline with the control defaulting off and indefinite retention |
Enforcement reaches the account; it does not reach the artifact#
Mali is the corpus's cleanest statement of a limit that applies to every disruption claim in this report:
"The platform ran fully on-premises using local models. The actor used Claude to provide software design and engineering support. Account enforcement actions do not affect the deployed product." "Our account enforcement actions disrupted the actor's software and design activities, but not the deployment of the platform."
A population-scale interception system for a state whose security services the US State Department and Human Rights Watch have documented detaining and abducting opposition figures, journalists and civil-society members, is running, on local models, and Anthropic's only available lever was the consultant's subscription. The conventional-weapons section supplies the same shape twice — the Yemen guided-weapons cell had already built "an offline simulation toolkit that does not rely on Claude," and the Russian drone-swarm team's code was written into their own project files.
The general form: the model's contribution is a durable artifact and the enforcement is a revocable account. Every "we disrupted the activity" in this report should be read against it. Disruption is measured at the vendor's boundary, and for the cases where the deliverable was software, that boundary is behind the harm.
Where the safeguards held, and where they did not#
The report is unusually direct, and the failures are more informative than the successes:
- "Our existing safeguards did not perform uniformly in these cases. In one case, Claude correctly refused a request but was overcome on further prompting. In another, it complied across many sessions without intervention." The re-prompt case produced "functional suppression guidance naming 10 private citizens to target for 'control'" across categories including petition interdiction and "talk to" coercive summonses.
- GTG-34007: "Claude refused explicit profiling and propaganda requests, but our safeguards did not refuse many of the surveillance software tooling requests." This is the load-bearing asymmetry. The refusal boundary tracks the stated act — profile this person — and not the function of the artifact — a phone-number-to-identity resolver, a messenger de-anonymizer, a front end for a case-management system with an 'action' tab. A tool request reads as engineering; the engineering is the surveillance.
- Declines did occur and were specific: covert interrogation and large-scale persona cultivation (GTG-14010), an ingest-and-produce weekly 'stability maintenance' report (GTG-14021, before the re-prompt).
Both failure modes are instances of Safeguard Evasion by Task Decomposition — the re-prompt reversal and the intent/artifact gap — reached here independently of the weapons and biological cases that page is mostly built from.
Connections#
- AI-Enabled Influence Operations — the sibling section, frequently the same actor: the CAR propaganda operation also ran a standing surveillance program on opposition figures, and the Iranian propaganda institutions and surveillance units share a state apparatus; production and monitoring are two ends of one program
- Safeguard Evasion by Task Decomposition — the two failure modes above, and the generalization: a refusal keyed to stated intent does not bind a request for a tool whose only use is the refused act
- The Stolen Model-Access Economy — how the Iranian units reached a blocked service, and the commercial layer: GTG-14010 also "drafted surveillance platform tenders and capability brochures marketed to bureau-level PRC government clients," a government-client-to-vendor structure that buys its AI the same way
- Claude Code — used with custom skills to drive a municipal cyber police sentiment pipeline against a government surveillance database; the harness's extension mechanism as surveillance tooling
- Agent Context Files — the institutional AI-usage manual codifying a prompt formula for bureau-wide distribution is the same artifact class, written by a state security office for its own staff
- AI-Accelerated Offense (hub) — the same labor-collapse thesis on the repression side: an analyst desk of many teams reduced to one office, a national interception platform built by one consultant
- Illicit Distillation — the collision between the two sections: a PLA-affiliated CCTV surveillance session and a municipal Public Security Bureau case-management build both reached Claude without their operators knowing, relayed by Moonshot and DeepSeek
- Google Threat Intelligence Group (GTIG) — the second provider's account: UNC5792's AI-classified Telegram monitoring and the exfiltrated-data-to-intelligence-report pattern in two espionage groups
- Anthropic — the author, the enforcement actor, and the party whose enforcement limit the Mali case documents
Open Questions#
- Mali, Yemen and the Russian drone swarm all leave a working artifact behind an account ban. Is there any case in the corpus where a vendor's enforcement demonstrably removed a deployed capability rather than the actor's continued access to the vendor?
- The refusal boundary tracks stated intent and not artifact function (a de-anonymizer, an identity resolver). Is there a published evaluation of refusal behavior on tool-construction requests whose only use is a prohibited act, separate from requests to perform the act?
- "Thousands of investigations per month" from a single office, and "6,388 Iranians profiled in a single year", are the only throughput figures in the section, and both are the actor's own claims repeated by the vendor. Does any external reporting corroborate the scale of AI-assisted state surveillance output?
Sources#
- Detecting and countering misuse of AI: September 2026 — Anthropic Threat Intelligence, Detecting and countering misuse of AI: September 2026, 2026-09-10,
case-study(first-party; disruption and safeguard-performance claims are the vendor grading itself, though the safeguard verdicts here run against its interest). The "Surveillance operations" section, pp. 81–110, plus GTG-30004/30005/30006 at pp. 103–110 (the GTG-34007 tooling asymmetry at p. 102, the Mali enforcement statements at p. 104, GTG-30006's refusal figure at p. 107): the three-trend framing, the eight case studies, the per-case key findings and attribution-confidence statements, the "our existing safeguards did not perform uniformly" paragraph (GTG-14021 disruption notes, p. 97), the GTG-34007 profiling-versus-tooling asymmetry, and the two Mali statements that account enforcement did not affect the deployed platform (p. 104). The case table above is assembled from case prose and key-findings bullets. Table note: the section's per-case workstream tables (GTG-14010, 14020, 14021, 14022, 50027) were read but no row is quoted that the surrounding prose does not restate — the ingesttable-weldwarning landed on wrapped row labels on p.91 ("Leadership of the Presbyterian Church in Taiwan"), verified as a single row againstpdftotext -layoutand not a merge - GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AI — Google Threat Intelligence Group, GTIG AI Threat Tracker: From Prompting to Autonomy, 2026-09-08,
case-study(first-party; no scale figures). Cited for UNC5792's Telegram monitoring bots and for the RAVINE CASTLE / CALANQUE ION exfiltrated-data summarization entries (prose and Tables 5–6)
Cited by 9
- AI-Enabled Influence Operations×2
The report says AI helped build the apparatus that "would otherwise need a staffed program office."…
- Anthropic×2
Detecting and countering misuse of AI: September 2026 is the fourth in a series (March, August and…
- Open Questions Backlog×2
Ai Enabled State Surveillance ×2 (oldest 12d) — Mali, Yemen and the Russian drone swarm all leave a…
- Agent Context Files
Ai Enabled State Surveillance — the same artifact class written by a state security bureau for its…
- AI-Accelerated Offense
Ai Enabled State Surveillance / Ai Enabled Influence Operations — the same labor-collapse thesis…
- Illicit Distillation
Ai Enabled State Surveillance — the collision between the two sections: a PLA-affiliated CCTV…
- Agent Security
Ai Enabled State Surveillance — Eight disrupted operations in Anthropic's September 2026 threat…
- Safeguard Evasion by Task Decomposition
Ai Enabled State Surveillance — the re-prompt reversal and the profiling-versus-tooling gap,…
- The Stolen Model-Access Economy
Ai Enabled State Surveillance — the commercial layer of the same market seen from the buyer's side:…
Related articles
- AI-Enabled Influence Operations
Nine disrupted campaigns in Anthropic's September 2026 threat report show the model building the *apparatus* of an infl…
- Autonomous Intrusion
The class of attack in which a model or a collective of agents conducts a network intrusion end-to-end — the campaign r…
- The Stolen Model-Access Economy
AI credentials have become loot, compute and cover at once — resale value, attack workloads run at the victim's expense…
- Agent Supply Chain Risk
Runtime-composed agent ecosystems expand the supply-chain attack surface: model poisoning (250 docs backdoor a 13B mode…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
