Sources#
- Claude Opus 5 System Card
- Security Incident INC-2026-07-28-01
- UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities
What it is#
The US Center for AI Standards and Innovation (CAISI) is the American government's frontier-AI evaluation body and the structural counterpart to the [[uk-ai-security-institute|UK AI Security Institute]]. In this corpus it always appears alongside AISI rather than alone — the two publish jointly, co-build benchmarks, and notify each other of incidents — which makes the pair, not either one, the corpus's unit of independent government evaluation.
It publishes through NIST: the July 2026 Kimi K3 assessment appeared simultaneously on
aisi.gov.uk/blog and as a NIST news item. Nothing in the corpus states its reporting line,
staffing, or statutory basis, so treat the NIST housing as the only institutional fact established
here.
What it does (in this corpus)#
- Joint capability assessment of an open-weight model before its weights shipped.
Preliminary Assessment of Kimi K3's Cyber Capabilities
(2026-07-23,
empirical) is a UK AISI / CAISI co-publication on Kimi K3, released three days before the weights. It is the corpus's only third-party dangerous-capability measurement of an open-weight model, and the only cyber comparison of PRC and US frontier models by anyone who builds neither. See Open-Weight Elicitation Irreversibility and LLM-Driven Vulnerability Research. - The cross-benchmark capability aggregation. Figure 2 of that assessment — the "Overall Cyber Capability" Elo chart against release date — carries the in-chart credit "Source: U.S. Center for AI Standards and Innovation." The method is described only as "inspired by Item Response Theory," aggregating tasks across multiple benchmarks onto one latent scale where a 400-point rise equals a 10× increase in the odds of solving tasks, with the methodology deferred to "prior published reports" not in this corpus. Its self-declared failure mode is the useful part: a model estimated from one benchmark gets a visibly wider confidence interval on the shared axis than models covered by several, so the axis is comparable while the precision is not. That is the honest version of the aggregation problem multi-axis measurement addresses by adding axes instead.
- Benchmark co-construction. Co-built the Gray Swan indirect-prompt-injection benchmark (28 scenarios, 1,130 high-transferability attacks) with UK AISI and model developers — the successor instrument after Agent Red Teaming saturated. See Agentic Prompt Injection.
- Incident counterparty. Notified by UK AISI on 3 August 2026 during INC-2026-07-28-01, alongside the affected model developers — evidence that the cross-institute channel carries incidents, not only publications.
The disclosure asymmetry worth recording#
CAISI's own figure is the corpus's sharpest instance of an evaluator disclosing less than the vendors it grades. Figure 2 individually names and plots ten PRC-lab models (DeepSeek R1 / R1-0528 / V3.1 / V4 Pro, Alibaba QwQ / Qwen3, Kimi K2 Thinking / K2.5 / K2.6 / K3, GLM 5.2) while the entire US side is an unlabeled aggregate trendline with a confidence band. The same choice runs through the prose: every US figure in the assessment — the 76.2% ladder score, 20 of 41 arbitrary-code-execution solves, step 28.5 of 32 — belongs to "the most cyber-capable U.S. models," a group never enumerated. The consequence is concrete rather than rhetorical: no US number in the document can be checked against any vendor's own published score, so the gap it reports is unattributable and unreplicable in the direction that matters most. See Compute-Controlled Benchmarking.
Connections#
- UK AI Security Institute — the counterpart body; co-author, co-benchmark-builder, and the organization that notified it of an incident. The pair is the corpus's independent-evaluation unit
- LLM-Driven Vulnerability Research — the joint assessment's ExploitBench milestone data is the corpus's only per-rung breakdown of where an open model's exploit chain stops
- Open-Weight Elicitation Irreversibility — CAISI co-supplies the one worked answer to that page's "who actually audits an open-weight release?" question
- Agentic Prompt Injection — co-built the Gray Swan IPI benchmark that replaced saturated ART
- Unsanctioned Action in Capability Evaluations — notified as a counterparty during AISI's self-disclosed incident
- Kimi (Moonshot AI) / GLM (Z.AI) — the two open-weight lines the joint assessment measures against an anonymous US aggregate
- Compute-Controlled Benchmarking — its comparator anonymity is a disclosure counter-exemplar from the one party with no product to protect
Sources#
- UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities — Preliminary Assessment of Kimi K3's Cyber Capabilities,
UK AISI / CAISI, 2026-07-23 (
empirical): joint authorship, the NIST mirror, the "Cyber Capability Trends" IRT methodology paragraph, and Figure 2's CAISI source credit and one-benchmark confidence-interval caveat - Claude Opus 5 System Card — the Gray Swan indirect-prompt-injection benchmark built with UK AISI, US CAISI and other developers
- Security Incident INC-2026-07-28-01 — UK AISI's incident notification list, 3 August 2026
Cited by 10
- AI-Accelerated Offense×2
Uk Ai Security Institute / Caisi — the evaluators who put a number on the open-weight floor, and…
- Capability-Gated Model Fallback×2
Uk Ai Security Institute / Caisi — the evaluators who have to switch this architecture off to…
- Compute-Controlled Benchmarking×2
Uk Ai Security Institute / Caisi — the independent evaluators who adopted "report capability…
- GLM (Z.AI)×2
Uk Ai Security Institute / Caisi — the government evaluators who graded GLM-5.2 as "the most…
- Kimi (Moonshot AI)×2
Uk Ai Security Institute / Caisi — the only assessors of K3 in this corpus with no product to sell,…
- Open-Weight Elicitation Irreversibility×2
Uk Ai Security Institute / Caisi — the audit that happened: joint, pre-release, black-box through…
- The Open-Weight Frontier Gap×2
A fourth axis, and the first one nobody with a product measured: cyber capability (July 2026).…
- UK AI Security Institute×2
Caisi — the US counterpart: co-author of the Kimi K3 assessment, source of its IRT-derived…
- LLM-Driven Vulnerability Research
Uk Ai Security Institute / Caisi — the joint evaluators; the first per-rung exploit-ladder…
- Entities — People, Orgs, Tools & Projects
Caisi — The US government's AI-evaluation body, publishing through NIST; UK AISI's counterpart and…
Related articles
- UK AI Security Institute
UK government AI-evaluation body (Science of Evaluation team); its July 2026 test-time-compute study is the first indep…
- Kimi (Moonshot AI)
Moonshot AI's open-weight Kimi line — K2.5/K2.6 as 1T-class MoEs already circulating in this corpus (Inkling's post-tra…
- Autonomous Intrusion
The corpus's first in-the-wild intrusion driven end-to-end by autonomous models — Hugging Face's July 2026 breach, re-a…
- Open-Weight Elicitation Irreversibility
A wiki-drawn synthesis of Brown and Gemma 4: if dangerous capability scales with inference budget, then an open-weight…
- The Open-Weight Frontier Gap
Arena Text, June 2026: the top closed model leads the best open model by 33 Elo and the best *dense* open model by 57;…
