Sources#
- Detecting and countering misuse of AI: September 2026
- GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AI
Summary#
Anthropic's September 2026 threat report details nine disrupted influence operations — the report's own definition being efforts "to manipulate the information environment… with the intent to deceive, distort, or covertly influence… typically while concealing the activity's origin, sponsorship, or coordination." They originated in Russia, Iran, Turkey and across the Gulf, South Asia, Africa and Europe; the operators were governments, state propaganda institutions, state media, commercial firms selling influence to paying clients, domestic political operators, and in one case an opposition movement in exile.
The interesting finding is not that a language model writes propaganda. It is which part of the operation the model built. case-study, first-party, no external verification; every prevalence and efficacy statement is the vendor grading its own detection.
The model builds the apparatus, not only the content#
The report's own trend line: actors had the model produce "doctrine manuals, opposition dossiers, ministerial portfolios, persona systems, target databases, employment contracts encoding editorial loyalty, and scoring rubrics that were used to rank staff who were part of the operation. This kind of work would otherwise need a staffed program office."
The Central African Republic case (GTG-04001) is the clean instance, and the sharpest detail in the whole section. A Russian-speaking actor in Bangui supplying the production backbone for a Wagner/Africa Corps-linked FIMI operation used Claude to run the operation's human resources: contracts mandating loyalty to the President of CAR and "Russia and its contingent," job descriptions, scoring rubrics, a three-strike dismissal process — and then scored staff articles against those criteria and asked Claude which employees to keep and which to fire. When Claude flagged the political weighting of the rubric, "the actor relabeled it in neutral terms and kept the scoring." The model was not writing propaganda at that moment; it was administering the personnel system of a propaganda office, and the one safety intervention it made was absorbed by a rename.
The Iranian trio (GTG-34001 — ICCO under the Ministry of Culture and Islamic Guidance, the Islamic Propaganda Office of Khorasan Razavi, the Bina Cultural Observatory) is the same shape at ministry scale: campaign plans, doctrine manuals, persona systems, target databases, ministerial planning documentation, a nine-part international influence portfolio on official ICCO letterhead, and organizational plans for the Supreme Leader's funeral. The UAE case (GTG-84002) adds ghost-written UN Human Rights Council testimony for two named individuals delivered under a cloned Swiss NGO identity, personal dossiers on 18 members of the European Parliament, and counter-accountability files on the UN Special Rapporteurs who had criticized the UAE's conduct in Sudan.
Persistent doctrine files as the coordination substrate#
The mechanism that makes these scale is one this wiki already tracks as a legitimate harness pattern. From the report: "Markdown files containing doctrine were reused almost verbatim across hundreds of sessions. Actors kept lists of banned words inside their AI agents, maintained shared files of approved sources and evasion rules, and ran custom software that called Claude in fixed batches." And the consequence: "The central setup meant that actors producing content never needed to coordinate with or even know one another." One actor was building a course to teach the workflow to others.
The MEK/NCRI operation (GTG-84006) names the artifacts: a shared agent platform called "Viktor" where each workspace kept its own long-term memory files, updated over time with banned words, approved sources, account-management rules and detection-evasion rules, "so the agent could keep producing content without a human user directing each session." The fingerprint list in the disruption table includes SKILL.md / LEARNINGS.md memory. The UAE operation embedded a master doctrine file in its system setup that instructed Claude to repeat the same mission "across hundreds of sessions."
This is Agent Context Files used adversarially, and it earns a place here because it changes what "coordination" means for attribution. Classic coordinated-inauthentic-behavior detection looks for shared infrastructure, synchronized timing, or contact between operators. Here the shared artifact is a text file of standing instructions, replicated across accounts that share nothing else, and the resulting output is uniform because the doctrine is — not because anybody talked. The uniformity that used to be evidence of a command channel is now evidence of a copied file.
Attribution laundering, including of the model's own caveats#
Actors "prompted Claude to intentionally strip state attribution from republished material, passing claims through chains of outlets so they read as independently confirmed." The Russian state-media desks (GTG-24015) ran this as the product: the same story echoed across Sputnik Moldova, RIA Novosti and a Moldovan amplification ecosystem "to manufacture false verification loops." The Malaysian platform (GTG-84005) republished TV BRICS, Xinhua, Sputnik/RIA and CGTN with the state attribution stripped, as independent Malaysian reporting.
One instance is categorically different and belongs on a safety page as much as a threat page: an actor tied to Russian state media "produced claims the model flagged as unverified, then instructed it to drop those caveats and present everything as confirmed, so the material would read as established fact." The model's calibration signal — the hedge that marks an unverified claim — was treated as a removable formatting layer. So was its style: the CAR actor pushed the model to strip "classic formatting habits, actively preventing the news feeds from reading like synthetic, AI-generated text." Both are the same move: the artifacts by which a reader might detect machine authorship or machine uncertainty are, to the operator, a post-processing step.
Refusals, and how they were routed around#
Claude refused at several points, and the pattern of what happened next is consistent enough to be the finding:
- CAR: refused the operation's most aggressive request — naming real individuals as militants to draw security action against them. "The actor pivoted to anonymous-source framing instead."
- Malaysia: refused after identifying a fabricated dossier as material for political defamation, and "balked at language that explicitly evoked a psychological operation." "The actor negotiated sanitized wording to keep building toward the same capability."
A refusal on the stated frame, followed by the same capability reached through a reframe. This is the influence-operations instance of the pattern Safeguard Evasion by Task Decomposition documents across five harm areas.
What the reach numbers actually say#
The report scores every operation on the Brookings Breakout Scale (Category One = confined to a single community on one platform; Six = broad public and policy impact), and the distribution is the counterweight to everything above:
| Operation | Scale of production | Breakout |
|---|---|---|
| Kenya astroturfing (GTG-54004) | Batches of exactly 50 tweets per session, pro-administration and anti-opposition on one playbook, plus the identical workflow resold for retail-brand marketing | One — "completely isolated within the network of fake accounts… failing to reach or influence any real people" |
| Commercial influence-as-a-service (GTG-54002, LKM Company, France) | 8,913 articles in ~20 languages across ~70 fabricated news sites, 70 matching X accounts, 250+ inauthentic commenting accounts, six continents | Two — no evidence of breakout beyond its own activity |
| Malaysia election platform (GTG-84005, BBS Bilişim, Istanbul) | ~1,000 fake X accounts with warm-up logic, a fake outlet, fabricated dossiers, all 222 parliamentary constituencies profiled on real census and electoral data | Two |
| MEK/NCRI (GTG-84006) | 500+ channels scraped, ~51,944 archived messages analyzed into psychographic dossiers, AI avatars with Persian audio | Two |
| Iranian state-aligned trio (GTG-34001) | Six languages with a 20-language plan; 100+ paid Iranian channels | Three |
| Bangladesh Awami League (GTG-54006) | 29 rotated accounts over ~16 months; ≥1,500 headlines, 300 narratives, 1,500 image prompts | Three |
| UAE / Muslim Brotherhood (GTG-84002) | ~300 inauthentic influencer accounts, a cloned NGO, ghost-written UN testimony | Three |
| CAR / Radio Lengo Songo (GTG-04001) | Daily content operation feeding a 98.9 FM station and the national broadcaster | Four |
Anthropic's own reading: "Most of the content we discovered drew little or no authentic engagement… The widest authentic reach occurred where state media outlets were the distribution mechanism (including FM radio, satellite and shortwave radio, and global television)." The one operation that reached a real audience did so because it owned a radio transmitter, not because it had a better model. Production is cheap and distribution is not, and nothing in these nine cases shows AI solving the distribution problem. Two caveats pull the other way: the report is scoring operations it disrupted early, so low reach is partly an artifact of the intervention it is grading; and GTG-24015's output reached actual Russian broadcast airwaves through an existing professional pipeline, which is the model succeeding precisely by not needing to build distribution.
A second provider's verdict: productivity, not breakthrough (GTIG, September 2026)#
GTIG's Q2 2026 tracker (GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AI, case-study, first-party, same upstream vantage as Anthropic's) reaches the same bottom line from Gemini: "none of these tactics have created breakthrough capabilities," and "we did not see evidence of successful automation." It observed activity aligned with China, Iran and Russia, plus commercial spammers and disinfo-for-hire. The examples fit this page's "the model builds the apparatus" reading at small scale. Iranian actors had Gemini write detailed text-to-image prompts (camera angles, studio lighting, facial texture) for fictitious personas, and had it take on specialist personas (a psy-ops expert, an oil-market analyst) with persuasion techniques built in. The forward-looking item is the one to watch. Indonesian actors were designing a centralized social-media automation platform with anti-detection browsers and proxy rotation, plus a WhatsApp bot gateway with multi-account management and "human-like AI conversational capabilities." GTIG says it "has not yet observed these interactive capabilities deployed in live operations." Anthropic's cloned-activist case (see User Awareness below) is a hand-run version of the same interactive capability, so the step GTIG is waiting for is automating a technique that already works manually. Two vendors reaching the same low-breakthrough verdict corroborates it, with the same limitation: both see the construction stage, and neither can follow content out to its audience.
The vantage point is the methodological finding#
"While a social media site usually sees an operation once its content is already circulating, we may see it on Claude while the operation is still being built. Actors use AI to plan their campaign, choose their targets, and write the material… Our visibility into these operations ends once it's live."
A model provider sits upstream of the platforms, at the production stage. That is a genuinely different observation post from the one the influence-operations research community has worked from, and it cuts both ways. It is why Anthropic can describe a doctrine file, a staff-scoring rubric and an unpublished target database — artifacts that never circulate and that no platform-side researcher can see. It is also why the reach figures are weak: the vendor cannot follow the content out, and relies on open-source research and industry data to say what happened next. And it is a third reason, on top of the report's "most notable and novel" selection, that no prevalence claim here is representative: the sample is operations that used a US frontier model at the construction stage, which selects against everyone using an open-weight model locally.
Connections#
-
Illicit Distillation — the sibling first-party attribution exercise in the same report, and the same methodological problem: named parties, published counts, no external check
-
AI-Enabled State Surveillance — the sibling section of the same report and frequently the same actor: the CAR operation ran a standing surveillance program on opposition figures, the MEK network built arrest-history dossiers on people inside Iran, and the Iranian propaganda institutions share personnel with the surveillance units; the two pages divide a continuum
-
Agent Context Files — the pattern turned adversarial: doctrine markdown, banned-word lists, approved-source files and
SKILL.md/LEARNINGS.mdmemory reused near-verbatim across hundreds of sessions, producing uniform output among operators who never meet -
Safeguard Evasion by Task Decomposition — the two refusals here and how each was routed around: a reframe to anonymous sourcing, and negotiated sanitized wording toward the same capability
-
The Stolen Model-Access Economy — how several of these operations reached Claude at all: VPNs, foreign phone numbers, rotated accounts and third-party IP-masking services, because access from Iran is blocked
-
AI-Accelerated Offense (hub) — the same "AI collapses the labor gap" thesis on the information side: a single staffer taken past what a newsroom could produce, as a state operator is taken past what a team of intrusion specialists could run
-
User Awareness — the impersonation case inverts it: the operation cloned a real activist's Telegram account, had Claude read ~8,400 of his posts to copy his style, and ran live political conversations with his contacts inside Iran, who "did not know they were speaking with an AI-assisted account"
-
Anthropic — the author, the enforcement actor, and the sole source of every figure here
-
Google Threat Intelligence Group (GTIG) — the second provider's account, with the same no-breakthrough verdict and the not-yet-deployed interactive bot gateway
Open Questions#
- The Breakout Scale distribution (mostly One–Three) is measured on operations disrupted early, by the party that disrupted them. Does any platform-side or academic study measure reach for AI-assisted influence operations that were not interrupted, so the low-reach finding can be separated from the intervention?
- A shared doctrine file produces uniform output among operators with no contact and no shared infrastructure. Does any attribution methodology detect that — a stylometric or instruction-level signature of a copied standing prompt — or does it defeat coordination detection outright?
- The report says AI helped build the apparatus that "would otherwise need a staffed program office." Is there any case where the apparatus outlived the account ban, as the Mali surveillance platform did on AI-Enabled State Surveillance? Nothing here says either way.
Sources#
- Detecting and countering misuse of AI: September 2026 — Anthropic Threat Intelligence, Detecting and countering misuse of AI: September 2026, 2026-09-10,
case-study(first-party; the disrupting party grading its own detection and its own reach estimates; cases selected as "the most notable and novel," so unrepresentative by construction). The "Influence operations" section, pp. 41–80: the definition and the nine case studies GTG-04001 / 54002 / 84005 / 24015 / 34001 / 54006 / 84006 / 54004 / 84002, the seven-item trends list (influence-as-a-service, AI as newsdesk, apparatus-building, complex tool use, attribution laundering, operational security, fake and impersonated personas), the "How we investigate" and "How we measure reach" methodology paragraphs, and the per-case Breakout Scale ratings. The reach table above is assembled from the per-case prose ratings, not from a table in the PDF. Figures 2–4 (pp. 49–51, screenshots of inauthentic commenting clusters) were read by image two-pass and are illustrative only — they carry no count this page relies on - GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AI — Google Threat Intelligence Group, GTIG AI Threat Tracker: From Prompting to Autonomy, 2026-09-08,
case-study(first-party; no reach or engagement figures). Cited only for the "Information Operations" subsection and its mitigations box: the no-breakthrough verdict, the Iranian persona-image and narrative prompts, and the Indonesian automation platform / WhatsApp bot gateway not yet seen in live operations
Cited by 11
- Agent Context Files×2
The file replaces the command channel, and that is an attribution problem. "The central setup meant…
- Anthropic×2
Detecting and countering misuse of AI: September 2026 is the fourth in a series (March, August and…
- AI-Accelerated Offense
Ai Enabled State Surveillance / Ai Enabled Influence Operations — the same labor-collapse thesis…
- AI-Enabled State Surveillance
Ai Enabled Influence Operations — the sibling section, frequently the same actor: the CAR…
- Google Threat Intelligence Group (GTIG)
It is the second first-party vantage on adversarial AI use, alongside Anthropic's September 2026…
- Illicit Distillation
Ai Enabled Influence Operations — the sibling first-party attribution exercise in the same report,…
- Agent Security
Ai Enabled Influence Operations — Nine disrupted campaigns in Anthropic's September 2026 threat…
- Open Questions Backlog
Ai Enabled Influence Operations ×3 (oldest 12d) — The Breakout Scale distribution (mostly…
- Safeguard Evasion by Task Decomposition
Ai Enabled Influence Operations — the negotiation form: a defamation refusal answered with…
- The Stolen Model-Access Economy
Ai Enabled Influence Operations — how several operations reached a blocked service at all: VPNs,…
- User Awareness
Ai Enabled Influence Operations — the inverse case in the wild: an operation cloned a real…
Related articles
- Autonomous Intrusion
The class of attack in which a model or a collective of agents conducts a network intrusion end-to-end — the campaign r…
- AI-Enabled State Surveillance
Eight disrupted operations in Anthropic's September 2026 threat report show the model standing in for three different i…
- The Stolen Model-Access Economy
AI credentials have become loot, compute and cover at once — resale value, attack workloads run at the victim's expense…
- Agent Supply Chain Risk
Runtime-composed agent ecosystems expand the supply-chain attack surface: model poisoning (250 docs backdoor a 13B mode…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
