資料來源#
摘要#
"From AGI to ASI" 報告以有效運算量為預測基礎——這是一個將三項彼此獨立改進的因素相乘而得的單一成長率。Epoch 估計其為每年約 10 倍(每年增加一個數量級),作者稱這是偏保守的下限數字。這是唯一擁有歷史資料、可用來套合預測模型的通往 ASI 的路徑,因此屬於「一切照常」的擴展,也是報告中最容易量化處理的抓手。
三項乘法因素#
| 因素 | 速率 | 備註 |
|---|---|---|
| 硬體製造(Moore's law 與相關因素) | 每年約 1.5 倍 | 每美元運算量,持續六十年——最不具不確定性的因素 |
| 運算投資成長 | 每年約 2.5 倍 | 過去十年硬體支出的成長 |
| 演算法效率 | 每年約 3 倍(Epoch:最高約 6 倍) | 達到固定效能門檻(例如 AlexNet-on-ImageNet)所需的 FLOPs,以約為 Moore's law 兩倍的速率下降;主要是許多增量式改進疊加,而非罕見的突破 |
硬體 × 投資 ≈ 每年 4 倍,代表投入最大型訓練運行的運算量。納入演算法效率(也就是「彷彿硬體機群成長」的效果)後,得到每年約 10 倍的有效運算量(1.5 × 2.5 × 3 ≈ 11.25,向下取整)。持續十年後,將比今天增加 10,000 倍。各因素的不確定性會彼此累積,因此真實速率可能顯著更高或更低——而且可能正在加速。
關鍵未解問題:運算是否會轉化為能力?#
運算成長是可以處理的預測對象;它如何轉化為新能力則不是。可能存在三種制度:報酬遞減(進展緩慢)、成比例(指數成長),或在遞迴式改進下呈超指數成長。報告的關鍵細節是:即使個別模型的進展停滯,持續增加運算量仍可透過執行更多實例、以更快速度、思考更久來提升整體能力。因此,「僅僅」量化擴展也可能解鎖看似質化的進步——例如一年內從 1,000 個 AGI 實例增加到 10,000 個,五年內增加到 1 億個(或以 100 倍速度執行 100 萬個實例)。這是否構成 ASI,正是擴展爭論的主軸(見 The Bitter Lesson:「更多運算 → 更多搜尋 → 更多智慧」;但要注意,天真的暴力搜尋在玩具領域之外會失效,真正的增益來自更好的先驗與啟發式方法)。
資料牆#
第一項重大摩擦是:高品質資料即將耗盡,無法再用來預訓練更大型的模型;預計本十年稍晚會開始造成影響(Villalobos et al. 2024)。模型規模的成長速度已超過新穎人類文字的產出速度。報告權衡的對策包括:
- 合成/自我生成資料——在天真的迭代訓練中有退化風險(Shumailov et al. 2024),但將測試時搜尋的輸出蒸餾回模型(AlphaZero 式)可以產生「略超越前沿」的資料;若數十億使用者持續消耗測試時運算量,這可能成為真正的遞迴式改進引擎。
- 模擬與互動資料(RL、多代理、生成式代理模型)——只要存在良好的模擬器,就能隨運算量直接擴展;例如 DeepMind 的 Adaptive Agent。
- 其他模態(影像/音訊/影片)可以延長跑道,但單靠人類生產無法以足夠速度成長。
結論:這很可能是摩擦,而非根本性阻礙——如果 ASI 由擴展驅動,資料生成有理由透過運算量以相近速度擴展。
經濟與資源摩擦#
如果進展主要依賴擴展,真正的限制問題是:跨越許多數量級的擴展,其經濟成本是否能夠持續——而這又循環取決於 AI 所產生的經濟報酬。相鄰限制包括:能源建設、土地/水資源、稀土,以及環境足跡(軌道資料中心等異想天開的提案也各自帶有風險)。即使擁有原始 FLOPs,記憶體頻寬與互連瓶頸仍可能限制有效利用率。反之,若進展來自演算法創新/自我改進/典範轉移,所需的經濟投入成長會較慢,而這只是一項邊際摩擦。
相關連結#
- AGI-to-ASI Pathways — 擴展是路徑 1;本頁是其量化引擎,以及兩項主要摩擦(資料牆、經濟)
- Intelligence Explosion Dynamics — 運算成長是遞迴迴圈加速的基底;報酬是成比例還是雙曲線,決定了所處制度
- Task Time-Horizon Scaling — METR 的時間範圍趨勢線,是這條運算側曲線在能力側的互補項(Whitfill et al. 模擬運算預測下的時間範圍成長)
- The Bitter Lesson — 「擴展是否足夠?」是作為預測問題的苦澀教訓;搜尋需要良好先驗,而不只是更多 FLOPs
- Multi-Agent Collective Intelligence — 「模型停滯但實例更多」的論點,將擴展導向集體能力
- Fundamental Limits of ASI — 為何能力預測必須以經驗為先:理論只能產生空泛的否定結論
- Advantages of Digital Intelligence — 這些正是會隨運算量擴展的 AI 特性,因此更多有效運算量會擴大人類與 AI 之間的差距
- Universal AI (AIXI) — AIXI 的近似版本保證會隨運算量改進,但暴力版本需要極其難以負擔的高速成長,才能帶來線性的智慧增益——這是「擴展是否足夠?」的理論後盾
- Inference Efficiency as Capability — 從推論側看演算法效率項:KV-cache、量化與 speculative-decoding 的增益,會在服務時而非僅在訓練時提高每美元的有效運算量
- Researcher Uplift from Code Output — 勞動與運算 R&D 分解中的運算側項目:Kwa 指出,每年運算量增加三倍,已可透過運算量約 0.45 的指數獨立地使研究投入每年增加約 1.6 倍,不涉及任何勞動增益,因此兩項投入都會複合成長
待解決的問題#
- 什麼時候更多運算能可靠地帶來更多智慧——只對某些問題類別如此,還是普遍如此?量化擴展與質化擴展能否互相權衡?
- 資料生成(合成、模擬、互動)真的能跟上模型規模的成長嗎?還是資料牆會先成為限制?
- 擴展何時(如果真的會)變得在經濟上不可行?硬體/軟體效率趨勢又會如何推移這個臨界點?
資料來源#
- From AGI to ASI — 第 2 節(有效運算量成長因素)、第 5.1 節(擴展路徑)、第 5.5 節(資料牆、經濟),表 4
Cited by 17
- AGI-to-ASI Pathways×3
Effective Compute Scaling — pathway 1's quantitative engine and its data-wall/economic frictions
- Open Questions Backlog×3
Effective Compute Scaling ×2 (oldest 58d) — When does more compute reliably yield more intelligence…
- Balance-of-Power Superintelligence×2
The numeric illustration — a self-improving system optimizing its own efficiency could "squeeze…
- Fundamental Limits of ASI×2
Effective Compute Scaling — why forecasting is empirical-first: theory gives only vacuous negatives
- Intelligence Explosion Dynamics×2
Effective Compute Scaling — exponential compute growth is the substrate a recursive loop bends…
- Multi-Agent Collective Intelligence×2
Effective Compute Scaling — "individual model plateaus but run more instances" routes scaling into…
- Researcher Uplift from Code Output×2
Effective Compute Scaling — the "compute tripling yearly" term in the R&D-speedup reconciliation is…
- RSI Growth Curves: Which Friction Binds First?×2
Data wall (Effective Compute Scaling) · (folded into "supply chain" in Future 1) · Demoted by both…
- Universal AI (AIXI)×2
This is the hub for the theory-of-superintelligence cluster: the formal anchor that Agi To Asi…
- Advantages of Digital Intelligence
Effective Compute Scaling — these advantages are precisely the ones that "scale with compute," so…
- Cross-Lab Pre-Release Review
Asked whether governments need to be involved, Musk volunteers that the group "probably should…
- Elon Musk
The compute geography — Effective Compute Scaling: power and cooling bind outside China, chips bind…
- Inference Efficiency as Capability
Effective Compute Scaling — efficiency gains enter the "effective compute" numerator the same way…
- Kimi (Moonshot AI)
He treats K3's efficiency as the headline, not its score. Chinese labs are "doing as well as they…
- Superintelligence Trajectory
Effective Compute Scaling — DeepMind's framing of compute growth as ~10×/year of 'effective…
- Task Time-Horizon Scaling
Effective Compute Scaling — the compute-side curve this capability-side trendline complements;…
- The Bitter Lesson
Effective Compute Scaling — "is scaling enough?" is the bitter lesson posed as a forecasting…
Related articles
- Recursive Self-Improvement
An AI system autonomously designing and developing its own successor; Anthropic Institute's *When AI builds itself* arg…
- Intelligence Explosion Dynamics
The growth-curve question behind recursive self-improvement: whether AI-accelerating-AI produces exponential, super-exp…
- The Abstraction Barrier
Lerchner's hypothesis that AI trained on human concepts may be unable to discover genuinely novel conceptual primitives…
- Artificial Superintelligence (ASI)
DeepMind's informal characterization of ASI as a system that exceeds large, well-coordinated human-expert collectives a…
- AGI-to-ASI Pathways
DeepMind's four non-exclusive, parallel technological routes from human-level AGI to superintelligence — scaling, algor…
