H
Howardism
Plate IIEntitiesHOWARDISM

US Center for AI Standards and Innovation (CAISI)

PublishedAugust 12, 2026FiledEntityDomainEntitiesTagsEntityOrgAI EvaluationGovernanceCybersecurityReading5 minSourceAI-synthesised

The US government's AI-evaluation body, publishing through NIST; UK AISI's counterpart and co-evaluator — joint author of the July 2026 Kimi K3 cyber assessment and the credited source of its cross-benchmark IRT/Elo capability analysis, co-builder of the Gray Swan indirect-prompt-injection benchmark, and a notification recipient in AISI's own INC-2026-07-28-01

Illustration for US Center for AI Standards and Innovation (CAISI)

Sources#

What it is#

The US Center for AI Standards and Innovation (CAISI) is the American government's frontier-AI evaluation body and the structural counterpart to the [[uk-ai-security-institute|UK AI Security Institute]]. In this corpus it always appears alongside AISI rather than alone — the two publish jointly, co-build benchmarks, and notify each other of incidents — which makes the pair, not either one, the corpus's unit of independent government evaluation.

It publishes through NIST: the July 2026 Kimi K3 assessment appeared simultaneously on aisi.gov.uk/blog and as a NIST news item. Nothing in the corpus states its reporting line, staffing, or statutory basis, so treat the NIST housing as the only institutional fact established here.

What it does (in this corpus)#

  • Joint capability assessment of an open-weight model before its weights shipped. Preliminary Assessment of Kimi K3's Cyber Capabilities (2026-07-23, empirical) is a UK AISI / CAISI co-publication on Kimi K3, released three days before the weights. It is the corpus's only third-party dangerous-capability measurement of an open-weight model, and the only cyber comparison of PRC and US frontier models by anyone who builds neither. See Open-Weight Elicitation Irreversibility and LLM-Driven Vulnerability Research.
  • The cross-benchmark capability aggregation. Figure 2 of that assessment — the "Overall Cyber Capability" Elo chart against release date — carries the in-chart credit "Source: U.S. Center for AI Standards and Innovation." The method is described only as "inspired by Item Response Theory," aggregating tasks across multiple benchmarks onto one latent scale where a 400-point rise equals a 10× increase in the odds of solving tasks, with the methodology deferred to "prior published reports" not in this corpus. Its self-declared failure mode is the useful part: a model estimated from one benchmark gets a visibly wider confidence interval on the shared axis than models covered by several, so the axis is comparable while the precision is not. That is the honest version of the aggregation problem multi-axis measurement addresses by adding axes instead.
  • Benchmark co-construction. Co-built the Gray Swan indirect-prompt-injection benchmark (28 scenarios, 1,130 high-transferability attacks) with UK AISI and model developers — the successor instrument after Agent Red Teaming saturated. See Agentic Prompt Injection.
  • Incident counterparty. Notified by UK AISI on 3 August 2026 during INC-2026-07-28-01, alongside the affected model developers — evidence that the cross-institute channel carries incidents, not only publications.

The disclosure asymmetry worth recording#

CAISI's own figure is the corpus's sharpest instance of an evaluator disclosing less than the vendors it grades. Figure 2 individually names and plots ten PRC-lab models (DeepSeek R1 / R1-0528 / V3.1 / V4 Pro, Alibaba QwQ / Qwen3, Kimi K2 Thinking / K2.5 / K2.6 / K3, GLM 5.2) while the entire US side is an unlabeled aggregate trendline with a confidence band. The same choice runs through the prose: every US figure in the assessment — the 76.2% ladder score, 20 of 41 arbitrary-code-execution solves, step 28.5 of 32 — belongs to "the most cyber-capable U.S. models," a group never enumerated. The consequence is concrete rather than rhetorical: no US number in the document can be checked against any vendor's own published score, so the gap it reports is unattributable and unreplicable in the direction that matters most. See Compute-Controlled Benchmarking.

Connections#

Sources#

§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 10
Related articles
  • UK AI Security Institute

    UK government AI-evaluation body (Science of Evaluation team); its July 2026 test-time-compute study is the first indep…

  • Kimi (Moonshot AI)

    Moonshot AI's open-weight Kimi line — K2.5/K2.6 as 1T-class MoEs already circulating in this corpus (Inkling's post-tra…

  • Autonomous Intrusion

    The corpus's first in-the-wild intrusion driven end-to-end by autonomous models — Hugging Face's July 2026 breach, re-a…

  • Open-Weight Elicitation Irreversibility

    A wiki-drawn synthesis of Brown and Gemma 4: if dangerous capability scales with inference budget, then an open-weight…

  • The Open-Weight Frontier Gap

    Arena Text, June 2026: the top closed model leads the best open model by 33 Elo and the best *dense* open model by 57;…