資料來源#
- Claude Fable 5 and Claude Mythos 5
- Claude Opus 4.8 System Card
- More compute, more capability: Why AI agent evaluations need to account for test-time compute
- Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown
摘要#
UK AI Security Institute(AISI) 是英國政府負責評估前沿 AI 能力的機構。其 Science of Evaluation 團隊 會在大型測試時運算預算下,於代理基準測試(網路、軟體工程、數學、學術、醫療保健)上執行前沿模型。由於它位於各模型實驗室之外,其數據可作為模型供應商所宣稱能力的獨立、政府研究機構佐證——這與 METR 對時間跨度曲線所扮演的角色相同。
其 2026 年 7 月的部落格文章 More compute, more capability 是本語料庫首篇 AISI 一手出版物,也是首個對大規模測試時運算論點群提供獨立 empirical 確認的出版物——截至目前,這組主張幾乎完全建立在一位 OpenAI 研究員(Noam Brown,practitioner-opinion)的研究之上。Brown 從軼聞出發主張「能力是預算的函數」;AISI 則在多個基準測試上實際量測了這一點。
在本語料庫中的職責#
- 測試時運算評估(2026 年 7 月)。 Science of Evaluation 團隊將 token 預算從低到高進行掃描,回報能力曲線,而非單一分數。研究結果顯示:約 8% 的網路任務只有在預算達到 ≥10M tokens 後才解出(部分任務高達 50M),在較小預算下完全看不見;最新模型在 100M+ 下仍持續上升;將預算從 1M 提高至 10M,使軟體工程分數提高約 25%(TerminalBench 2.0、SWE-Bench Pro),數學/學術分數提高約 22%(Humanity's Last Exam)。其 2026 年 3 月的前身研究首次指出,適度的運算上限會低估網路能力——這正是 Brown 引用的結果(模型「在 100M tokens 下仍持續改善」)。
- 運算需求—人類時間法則。 在其網路任務與 METR 的軟體工程任務中,代理所需的運算量會隨熟練人類完成任務所需的時間而擴張——這是一條擬合指數約為 0.7–1.0 的冪律(一分鐘任務約需數千 tokens、一小時約需數百萬、一週約需數十億)。
- 網路 CTF 測試套件。 維護一套狹義的網路 capture-the-flag 任務(Fig-4 分析中有 78 項),包括約需 20 個人類工時的 「The Last Ones」——受測模型中沒有任何一個能在低於 30M-token 預算下完成。並將 METR 的 211 項軟體工程任務集與自有任務一併使用。
- Agent Red Teaming(ART)。 共同維護 Gray Swan/UK-AISI ART 基準測試;Claude 模型大多已在其中達到飽和(見代理式提示注入)。
- 模型紅隊測試。 在簡介 5上,這是唯一被指出曾在短暫初始窗口內朝通用 jailbreak取得部分進展的紅隊測試組織;其他外部紅隊測試者則毫無所獲(見能力門控模型回退、由 LLM 驅動的漏洞研究)。
為何在此重要#
AISI 的曲線將測試時運算論點從實驗室研究者的 practitioner-opinion 推進為經量測、獨立重現的事實,也把抽象的「回報預算」規範轉化為實務改變:現在它會跨多個預算進行評估(對最困難的任務也包括非常大的預算),依據預算回報可靠性與可達範圍,避免資源不足的評估被誤認為低能力模型,並正在定義**「最低資訊量預算」(只有當增加更多運算後可達範圍不再上升,才宣告已達到上限)。它也將以較便宜的執行結果預測高預算效能**列為明確且尚未解決的研究方向,並正積極投入——這正是 Brown 僅提出的開放問題。
相關連結#
- 大規模測試時運算 — 以實證佐證核心論點;Brown 引用的 AISI 網路評估是 AISI 自己的工作
- 運算控制基準測試 — 「回報能力曲線」是政府評估者對「將運算量放在 x 軸上」的具體實現
- 任務時間跨度擴張 — 顯示時間跨度及其倍增速率都取決於預算;重新使用 METR 的任務集
- 潛在能力過剩 — 量測能力過剩:約 8% 的網路任務在低於 10M tokens 時不可見,「The Last Ones」則低於 30M 時不可見
- 負責任擴展政策評估 — 在自身的安全評估實務中,將對 RSP/準備度框架之無上限預算批評予以操作化
- 開放權重能力引出不可逆性 — 其實證曲線為「危險能力會隨預算擴張」這項前提提供量測上的支持
- 代理式提示注入 — 共同維護 ART 代理紅隊測試基準
- 能力門控模型回退 / 由 LLM 驅動的漏洞研究 / Claude Fable 5 — 其在 Fable 5 上取得部分通用 jailbreak 進展
- METR — 同級的獨立第三方評估者;AISI 在運算需求分析中重新使用其 211 項軟體工程任務集
- Noam Brown — 其測試時運算論點獲 AISI 獨立佐證的 OpenAI 研究員
資料來源#
- More compute, more capability: Why AI agent evaluations need to account for test-time compute — More compute, more capability(2026-07-02,
empirical):能力曲線、網路 CTF 預算、運算需求—人類時間冪律、取決於預算的時間跨度,以及三個開放研究問題 - Claude Fable 5 and Claude Mythos 5 — AISI 在 Fable 5 上取得部分通用 jailbreak 進展
- Claude Opus 4.8 System Card — Claude 模型已在其中達到飽和的 Gray Swan/UK-AISI Agent Red Teaming(ART)基準測試
- Really Big Test-Time Compute in AI Changes Benchmarks, Safety and Research with OpenAI's Noam Brown — Brown 引用 AISI 的網路評估,指出模型在 100M tokens 下仍持續改善
Cited by 17
- Unsanctioned Action in Capability Evaluations×5
This lands awkwardly on AISI's own most-cited contribution to this wiki. Its July 2026 study is the…
- Compute-Controlled Benchmarking×4
Does a compute-controlled evaluation regime advantage frontier labs (who can afford the full curve)…
- Evaluation Awareness & Grader Gaming×4
Two things follow for this page. First, the confound is not marginal: on the alignment evals where…
- Large-Scale Test-Time Compute×4
Can high-budget performance be predicted from low-budget runs? Brown's proposed research question:…
- Task Time-Horizon Scaling×4
The UK AI Security Institute's July 2026 study (empirical) adds a confound the doubling curve above…
- Latent Capability Overhang×3
Uk Ai Security Institute — the government evaluator that measured the overhang: ~8% of cyber tasks…
- Open-Weight Elicitation Irreversibility×3
Dangerous capability scales with inference budget. Brown (practitioner-opinion): if a model "keeps…
- AI-to-AI Coercion×2
Six frontier managers, 30 conversations per cell (10 scenarios × 3 seeds), up to 12 manager turns…
- METR×2
Uk Ai Security Institute — sibling independent evaluator that reuses METR's task set and shows the…
- Noam Brown×2
Brown's essay is practitioner-opinion — arguments and anecdotes from one lab. In July 2026 the UK…
- Responsible Scaling Policy Evaluations×2
Brown's critique is now empirically demonstrated — by a government evaluator. The UK AI Security…
- Anthropic
It calls for the practice to spread: "We encourage other AI labs to perform similar reviews." UK…
- Automated Behavioral Audit
Petri trades depth for portability, and the card is explicit about the cost: about a quarter as…
- Benchmark Score Redundancy
Uk Ai Security Institute — the government evaluator pursuing the compute-axis version of this idea…
- Claude Opus 5
Uk Ai Security Institute — external cyber-range and misalignment testing
- Expenditure Horizon
The budget is unspecified. Time horizon "doesn't fully specify a budget or constraints for tokens…
- Entities — People, Orgs, Tools & Projects
Uk Ai Security Institute — UK government AI-evaluation body (Science of Evaluation team); its July…
Related articles
- Responsible Scaling Policy Evaluations
Anthropic's RSP gates deployment on pre-release capability evaluations in CBRN, automated AI R&D, and high-stakes misal…
- Compute-Controlled Benchmarking
Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…
- Claude Opus 5
Anthropic's Opus-class release of July 2026; matches Mythos 5 on capability without advancing the frontier, is the best…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Large-Scale Test-Time Compute
Noam Brown's thesis that model capability is now a function of inference budget (tokens/cost/time): with good scaffoldi…
