Sources#
Summary#
Claude Opus 5.5 is Anthropic's Opus-class model following Opus 5. The wiki's only source on it so far is not a system card but METR's third-party predeployment AI R&D evaluation (2026-09-22), conducted under an unpaid agreement, drafted by METR, then reviewed and edited by Anthropic before both parties approved the final text for incorporation into the Opus 5.5 system card. That provenance — the evaluator's own words edited by the evaluated party before publication — is a sharper form of the review-access conflict already tracked on METR, and it bounds every claim below: this is jointly-authored text about a model's own creator, not an independent report.
What METR found#
Capability testing ran over 10 business days via API access, across five tasks: Budget NanoGPT Speedrun (a constrained version of the NanoGPT Speedrun competition), Language Model Conceptual Argumentation (LMCA) (a conceptual-reasoning dataset, Cooper et al. 2026), Train a Program (training ML models to replicate software behavior), Gaming Bot (Python agentic play via a non-visual API), and Sunlight (open-ended research ability and report writing). METR supplemented this with background trends from its own Frontier Risk Report, an Anthropic capability/control questionnaire, and an Anthropic researcher interview.
(A) Acceleration capability of Opus 5.5 itself. "Acceleration from this model would be slightly higher than for Fable 5.1, but this model is unlikely to be able to fully automate AI R&D." Opus 5.5 is "a modest rather than huge leap" above Fable 5.1, with improvement shown on both verifiable tasks (Budget NanoGPT, Gaming Bot) and harder-to-verify ones (LMCA, Sunlight); Anthropic claims it continues the Mythos-level trend on its internal AECI metric. Full automation would require "large improvements in foresight, prediction, creating one's own feedback loops" and researcher judgment — and METR's evidence does not suggest Opus 5.5 is a major improvement on these. This is the same judgment/taste gap — see Research Taste as the Human Bottleneck — that every prior Anthropic model generation has hit. Still, "Claude Opus 5.5 is still likely to noticeably accelerate researchers and automate limited aspects of R&D," at a slightly higher uplift than Fable 5.1. METR is explicit that its data cannot distinguish a consistent, accelerating, or decelerating rate of AI-R&D improvement — only that frequent incremental gains are consistent with any of the three.
(B) AI's acceleration of Opus 5.5's own development. "The development of this model was at least somewhat accelerated by AI but is unlikely to have been dramatically accelerated by AI." The evidentiary weight here sits outside METR's own capability testing: a separate METR team, given elevated access inside Anthropic, ran an experimental assessment of internal AI-R&D acceleration and shared its conclusions with the publishing team — but "could not disclose supporting evidence or reasoning details." The publishing team states plainly that it "use[s] the evidence in the provided AI R&D report as an input to our assessment but do not argue directly in defense of its claims." That internal report's headline figure: "~1.5X overall acceleration in capabilities due to AI (i.e. 1.5 years in 1 year), with perhaps 30% chance of 2X acceleration" — timeframe unspecified. Opus 5.5 reads as an on-trend improvement over Fable 5.1 on Anthropic's internal capability evaluations, consistent with continuation of the AECI trend from Mythos Preview onward.
This is described as "a trial version of a more holistic assessment, with further public outputs expected" — an acknowledged first pass, not a settled instrument.
What the assessment does not cover#
METR states its own scope explicitly: the evaluation gathers AI R&D capability evidence and is "not meant to verify claims about compliance with any specific threshold from Anthropic's policies," and it does not assess alignment properties at all.
Connections#
- METR — the evaluator; this is METR's first predeployment AI R&D assessment in the corpus (distinct from its post-incident investigations), and the first case where Anthropic edited the evaluator's own text before publication
- AI R&D Autonomy Evaluation (AECI) — the AECI/AI-R&D-threshold framework this evaluation sits inside; Anthropic's claim that Opus 5.5 continues the AECI trend is unaudited here (no point estimate is given, unlike the Opus 5 card's 162.1)
- Research Taste as the Human Bottleneck — "foresight, prediction, creating one's own feedback loops" and researcher judgment are this page's residue, restated by a third party as the reason Opus 5.5 does not close the automation gap
- Claude Opus 5 — predecessor; no capability table exists yet for Opus 5.5 to compare against it directly
- Researcher Uplift from Code Output — the ~1.5X (30% chance of 2X) figure is a new, still-opaque data point on the same acceleration question Kwa's production-function estimate and Anthropic's Risk Report both address
Open Questions#
- What supporting evidence backs the internal AI R&D report's "~1.5X, perhaps 30% chance of 2X" acceleration estimate? The team that produced it had elevated access but did not disclose its reasoning even to the METR team publishing this assessment. The redacted-leading-indicators problem this raises is the same one tracked on AI R&D Autonomy Evaluation (AECI).
- Does a full Claude Opus 5.5 system card change the capability picture METR's five-task summary sketches — in particular, does an AECI point estimate exist, and how does it compare to Opus 5's 162.1?
Cited by 6
- AI R&D Autonomy Evaluation (AECI)×4
Every determination above is Anthropic assessing Anthropic, including the Risk Report's own AI R&D…
- METR×4
Claude Opus 5 5 — the September 2026 predeployment AI R&D evaluation, drafted by METR and then…
- Research Taste as the Human Bottleneck×4
METR's predeployment evaluation of Claude Opus 5.5 (empirical) reaches for the same vocabulary from…
- Claude Opus 5
Claude Opus 5 5 — direct successor; METR's predeployment evaluation reads it as a modest gain over…
- Entities — People, Orgs, Tools & Projects
Claude Opus 5 5 — Anthropic's Opus-class release following Opus 5; METR's predeployment AI R&D…
- Open Questions Backlog
Claude Opus 5 5 ×2 (oldest 5d) — What supporting evidence backs the internal AI R&D report's…
Related articles
- AI Accelerating AI Development
The empirical core of *When AI builds itself*: measured evidence AI already speeds AI R&D at Anthropic — >80% of merged…
- Researcher Uplift from Code Output
Thomas Kwa (METR) translates Anthropic's reported 8× code-per-engineer-per-day into serial researcher uplift with produ…
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- Structured Safety Case (Claim Decomposition)
Anthropic's August 2026 Risk Report replaces prose risk assessment with an explicit argument: misalignment risk decompo…
- AI R&D Autonomy Evaluation (AECI)
How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…
