資料來源#
摘要#
METR(Model Evaluation & Threat Research)是一個評估前沿 AI 能力的獨立組織,最廣為人知的是其 時間範圍測量:模型能可靠地獨立完成任務的長度。其資料是 Anthropic Institute〈When AI builds itself〉文章所依據的外部基準核心,也構成了本 Wiki 任務時間範圍擴展頁面的基礎。
它的工作#
- 時間範圍。 報告模型在一組任務中達到 50% 可靠度時所對應的任務持續時間(在 80% 可靠度下也呈現相同趨勢)。METR 的主要發現是,這個範圍大約每四個月翻倍,高於早期約七個月翻倍的速度——這是能力正在加速而不只是提升的量化證據。
- 前沿模型的長任務測量。 METR 發現,Claude Mythos Preview 能工作「至少」16 小時,並且「已達 [METR] 在不新增任務的情況下所能測量的上限」——也就是說,前沿模型已開始超越基準測試自身的天花板。
- 獨立的第三方訊號。 由於 METR 位於各實驗室之外,其數據可作為 Anthropic 內部加速主張的外部佐證,例如約 8 倍的程式碼產出量(AI 加速 AI 開發)。
- 被其他評估者重複使用。 UK AI Security Institute 於 2026 年 7 月進行的 test-time-compute 研究,使用 METR 的 211 項軟體工程任務集(以及 AISI 自有的網路安全任務),並延伸了時間範圍的框架:研究顯示,時間範圍及其翻倍速度都取決於預算(見任務時間範圍擴展)。
相關連結#
- 任務時間範圍擴展 — 建立在 METR 時間範圍指標上的概念頁面
- AI 加速 AI 開發 — METR 的外部趨勢線為 Anthropic 的內部產出量證據提供佐證
- 遞迴自我改進 — 將翻倍曲線外推後,它成為 RSI 比預期更早到來的量化依據
- Mythos Model — METR 評估為「至少 16 小時」、超出其目前測量上限的模型
- UK AI Security Institute — 重複使用 METR 任務集,並顯示時間範圍指標取決於預算的姊妹獨立評估者
- Researcher Uplift from Code Output — METR 的 Thomas Kwa 於 2026 年 7 月撰寫的建模備註,將 Anthropic 的 8 倍程式碼數據轉換為約 2.5 倍的序列研究人員提升;其關於冗長程度及主觀感受與實際加速差異的注意事項,依據的是 METR 自身的 uplift RCT
待解決的問題#
- 當目前的任務組合飽和後,METR 將建立哪些新任務,以測量持續數天及數週的時間範圍?
- METR 也執行了顯示開發者對 AI 提升效果的自我估計過高的研究——它要如何將這種懷疑態度,與自身陡峭的時間範圍曲線調和?進一步釐清: Researcher Uplift from Code Output — 一位 METR 建模者(Kwa)正好處理了這個矛盾:他折減自我報告(引用 METR 關於主觀感受為 +20%、實際結果為 −20% 的發現),並指出冗長程度的影響,但仍從一個客觀的 8 倍程式碼產出量數據,而非自我估計,推算研究人員提升幅度超過 2 倍——也就是說,METR 懷疑的是自我報告指標,而不是加速現象本身確實存在。
資料來源#
- When AI builds itself — 引用 METR 的時間範圍,以及 METR 對 Mythos Preview「16 小時/已達我們可測量範圍上限」的評估
Cited by 8
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Entities — People, Orgs, Tools & Projects
Map of Content for all 55 entity pages. See Home for concept domains.
- Mythos Model
Anthropic preview-tier frontier model and the first member of the Mythos-class tier (above Opus); gated for safety, use…
- Open Questions Backlog
_396 actionable open questions across 155 pages · 79 predictions · 9 notes · 21 in progress · 59 watching (entities), a…
- Researcher Uplift from Code Output
Thomas Kwa (METR) translates Anthropic's reported 8× code-per-engineer-per-day into serial researcher uplift with produ…
- Returns to Expertise in Agentic Coding
Anthropic's 400K-session study: domain expertise (not coding skill) is what amplifies an agent — experts get 2× the act…
- Task Time-Horizon Scaling
METR's measure of the task length AI can complete reliably on its own, doubling roughly every 4 months (up from every 7…
- UK AI Security Institute
UK government AI-evaluation body (Science of Evaluation team); its July 2026 test-time-compute study is the first indep…
Related articles
- AI R&D Autonomy Evaluation (AECI)
How Anthropic measures whether a model can automate or dramatically accelerate AI research — the capability that drives…
- AI Accelerating AI Development
The empirical core of *When AI builds itself*: measured evidence AI already speeds AI R&D at Anthropic — >80% of merged…
- Anthropic
AI safety company / vendor of Claude; mission-as-tiebreaker culture; ~30–40 PMs across teams; Mike Krieger leads Labs r…
- Claude Fable 5
Anthropic's first generally-available Mythos-class model (June 2026) — state-of-the-art on nearly all benchmarks; the s…
- LLM-Driven Vulnerability Research
Claude Mythos Preview's emergent cybersecurity capabilities: autonomous zero-day discovery, full exploit chains, and An…
