資料來源#
這是什麼#
Google DeepMind 的第四代開放權重模型家族,於 2026-07-02 以 Apache 2.0 授權發布(arXiv 2607.02770)。原生支援多模態(文字、影像、音訊),明確以「各式硬體環境」與「邊緣部署」為目標,而非追求前沿能力。
共有五款模型:
| 模型 | 參數量 | 備註 |
|---|---|---|
| E2B | 有效 2.3B(總計 5B) | 逐層嵌入,150M 視覺編碼器 + 305M 音訊編碼器 |
| E4B | 有效 4.5B(總計 8B) | 逐層嵌入,使用相同編碼器 |
| 12B | 12B 稠密 | 無編碼器 — 沒有視覺或音訊編碼器 |
| 26B-A4B | 總計 26B / 啟用 3.8B | Mixture-of-Experts,550M 視覺編碼器 |
| 31B | 31B 稠密 | 550M 視覺編碼器,旗艦款 |
僅解碼器 Transformer,前置正規化 + 後置正規化 RMSNorm、QKNorm。詞彙表含 262k 個 SentencePiece 詞元。預訓練資料截止於 2025 年 1 月 — 距發布已有十八個月。使用 TPU v5p/v6e(4,096–12,288 個晶片)訓練,搭配 Slice-Granularity Elasticity,將局部晶片故障造成的停滯時間「從許多分鐘縮短到幾秒」。
真正的新意#
大致依重要性由高至低,有四項。
開放權重模型加入思考模式。 Gemma 4 會先輸出推理軌跡,再給出回應;在開頭的 system turn 中放入 <|think|> 控制詞元即可啟用。這代表 test-time-compute 這條路線進入任何人都能下載和執行的模型家族。值得注意的是,任何人都能切換這個詞元 — 請見 Open-Weight Elicitation Irreversibility。
12B 捨棄了編碼器。 視覺方面:將 48×48×3 RGB patches 輸入單一 35M 矩陣乘法,搭配 2D 座標位置嵌入,取代 550M ViT。音訊方面:305M USM conformer 完全捨棄;原始 16 kHz 音訊切成 40 ms 區塊(640 維向量),直接投影至 LLM 嵌入空間。音訊本身已是時間序列,因此完全不使用位置編碼。這是另一個實驗室以不同理由印證 Encoder-Free Early Fusion。
深度推論效率技術組合。 KV 快取減少 37.5%,量化後縮至次 GB,並發布一個推測解碼起草器頭。詳見 Inference Efficiency as Capability — 這正是本報告值得收錄於 wiki 的原因。(2026-09-23:這 37.5% 的其中一項技術,現在有獨立研究結果提出反證。Gao 與 Xu 的 Keyless Attention 論文(arXiv 2606.21848,empirical)對 Gemma 4 採用的 values = keys 家族進行消融 — 即深度為 2、沒有替代路由矩陣的「KV-sharing」變體 — 在架構匹配的 12 層 GPT-2 上發現,其最佳驗證損失劣於 QKV 基線;而明確取代已刪除投影路由功能的深度為 3 變體則與基線相當。Gemma 4 報告這項技術沒有造成損失,卻未列出架構匹配的消融,因此這項主張如今仍屬 vendor-claim,且在小規模實驗中遭到一項 empirical 結果反駁。)
它無法與前沿模型競爭,而且坦然承認。 在 Arena Text(2026 年 6 月 19 日)中,Gemma 4 31B 排名 第 43,Elo 1451 ±8 — 是領先的稠密開放模型,比排名第 1 的 Claude Fable 5 低 57 Elo,並落後於六款規模大它 20–50 倍的 MoE 開放模型。請見 The Open-Weight Frontier Gap。
仔細解讀基準測試表格#
報告的重點比較(表 5)拿思考模式下的 Gemma 4,對比非思考模式的 Gemma 3 27B;論文中完全沒有同一款 Gemma 4 模型在思考與非思考模式下的消融比較。因此,在引用最多的表格中,世代差異與推論預算差異混雜在一起 — 這正是 Compute-Controlled Benchmarking 所描述的典型問題。長上下文表格(表 9)兩邊都未啟用思考模式,是文件中唯一乾淨的世代比較。
檢視論文本身文字的三項主張:
- **「E2B 以少 10 倍參數量,大致匹配 Gemma 3 27B」**呈現的是參差不齊,而非全面持平。E2B 在 AIME 2026(37.5 vs 20.8)、Codeforces Elo(633 vs 110)、LiveCodeBench v6(44.0 vs 29.1)勝出;但在兩項廣泛知識基準 — MMLU Pro(60.0 vs 67.6)和 MMMLU(67.4 vs 70.7)— 以及 τ²-airline(31.0 vs 39.0)上都落敗。推理能力可以壓縮,儲存的知識則不行。請見 Jagged Intelligence (Ghosts, Not Animals)。
- **「E4B 在所有[視覺]評測中與 Gemma 3 27B 相當或更好」**大致成立,InfographicVQA 除外(70.0 vs 70.6)。
- 裁掉視覺詞元時,無編碼器 12B 的表現隨規模變化而不成比例地退化。 將視覺詞元從 1120 減至 280,12B 在 InfographicVQA 上下降 −29.7 分,比 31B(−9.2)、26B-A4B(−11.5),甚至更小的 E4B(−15.2)和 E2B(−19.3)都跌得更多。OmniDocBench 1.5 也呈現相同反轉(12B 誤差 +0.244,是全系列最差)。這個異常只出現在影像中的密集文字任務;MMMU Pro 和 MATH-Vision 的退化則符合預期。論文未提及此事。相關記錄見 Encoder-Free Early Fusion — 截至 2026-09-21,另一個實驗室的同規模消融顯示,僅投影的視覺路徑在 OCRBench 上表現領先,因此問題更可能出在詞元預算,而非投影本身。
這個家族底部的能力落差很大。Humanity's Last Exam:31B 得分 19.5、26B-A4B 得分 8.7、12B 得分 5.2,兩款小模型則未公布數據。GraphWalks F1:E2B 得分 4.1,Gemma 3 27B 則為 32.8。
安全性:文件中的證據等級有所下降#
整份報告的證據標示為 evidence: empirical — 包含十六個實測結果表格。第 5 節則一張表都沒有。安全性主張都以文字敘述:「我們看到所有內容安全類別都有重大進步」、「Gemma 4 模型在提升安全性的同時,顯著優於 Gemma 3 和 3n 模型,且不合理拒答維持低檔」、「模型產生的政策違規極少」。沒有數字、沒有基準名稱,也沒有運算預算。
即使周圍的能力主張都有實證支持,這些安全性主張仍應視為 vendor-claim。在此特別重要的是,權重公開且永久可用 — 請見 Open-Weight Elicitation Irreversibility。
測量工具其實存在 — 只是拿去測封閉模型系列。 十九天後,DeepMind 發布了 Gemini 3.5 Flash-Lite 模型卡(2026-07-21),其中安全性章節以五列表格,列出相較前一款模型的百分點差異:文字轉文字 −8.14pp、多語言 −0.92pp、影像轉文字 0pp、語氣 +3.04pp,以及不合理拒答 +5.32pp,方向變差。最後一列正是 Gemma 4 第 5 節聲稱但未測量的指標(「不合理拒答維持低檔」)的實測版本;對象是同一實驗室、同一個月推出的姊妹模型 — 而當實驗室真的測量時,數字顯示情況朝不利方向變化。這並未推翻 Gemma 4 的主張(模型不同、系列不同、沒有共用基準),但排除了對缺少表格最有利的解釋。
相關連結#
- Inference Efficiency as Capability — 報告真正的貢獻:以五種方法降低單位推論成本
- Encoder-Free Early Fusion — Gemma 4 12B 是此設計的第二個獨立案例,也是第一個開放權重規模的案例
- The Open-Weight Frontier Gap — Gemma 4 對比開放 MoE 巨型模型與封閉前沿模型時的實際位置
- Open-Weight Elicitation Irreversibility — 思考模式、公開權重與未以表格呈現的安全性評測
- Compute-Controlled Benchmarking — 表 5 是混淆因素的實例
- Jagged Intelligence (Ghosts, Not Animals) — 「參數量少 10 倍」的主張呈現參差表現:推理能力能壓縮,知識則不行
- Google DeepMind — 所屬實驗室;Gemma 是其開放權重系列,與 Gemini 不同
- Large-Scale Test-Time Compute — 思考模式讓這條路線進入開放權重模型
- The Bitter Lesson — 報告一方面遵循其精神(捨棄編碼器),另一方面又違背其精神(手工設計推論路徑)
- GLM (Z.AI) — 另一個 2026 年開放權重模型家族,採取相反策略:以 750B-A40B 追求前沿能力,而 Gemma 以 31B 追求效率
待解決的問題#
- 為何 MoE 表現不如稠密模型?Gemma 4 26B-A4B 在 Arena 的 Elo 為 1438,低於 31B 的 1451,儘管他們表格中每款更大型的開放模型都採用 MoE 架構。論文未探討此事。
- 預訓練資料截止於 2025 年 1 月,但模型在 AIME 2026 得分 89.2。報告表示資料經過篩選以「去除基準污染」。對於一場在資料截止後才舉辦的競賽,這究竟意味著什麼?
- 無編碼器 12B 在密集文字任務上的退化,究竟是 35M 投影沒有進行特徵壓縮所致,還是 12B 特定訓練過程的結果?相同規模的有編碼器/無編碼器消融即可釐清;論文沒有進行。部分解答(2026-09-21): 另一個實驗室現在已做過一項消融,但僅限視覺 — Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation 使用相同的 Qwen2.5-7B-Instruct 主幹與相同資料,分別訓練只用 patch embedding 的 7B 模型,以及帶有 SigLIP-2 的 7B 模型;在每一個密集文字項目上,無編碼器版本都領先(OCRBench 79.7 vs 78.3、AI2D 79.6 vs 79.4、ChartQA 同為 85.6)。因此,僅投影的視覺路徑並非天生不擅長處理密集文字。但問題仍未解決,因為該消融只採用一種視覺詞元預算,而 Gemma 4 的 −29.7 分只在 280 個詞元時出現;此外,該評測套件不含 InfographicVQA 或 OmniDocBench。更有力的反證測試現在應該是掃描不同詞元數,而非使用更大的模型。
資料來源#
- Gemma 4 Technical Report — Gemma 4 Technical Report,Gemma Team、Google DeepMind(arXiv 2607.02770,2026-07-02)。能力部分標示為
empirical;第 5 節的安全性主張沒有表格佐證,因此視為vendor-claim。
Cited by 22
- Google DeepMind×6
This places DeepMind opposite Anthropic's Economic Index on the wiki's usage-measurement axis, and…
- Inference Efficiency as Capability×4
Right on the substrate. Gemma 4 (empirical, mid-2026) is the trend continuing exactly as predicted…
- Inkling×4
Attention: sliding-window and global layers interleaved 5:1 with 8 KV heads — the same ratio Gemma…
- Encoder-Free Early Fusion×3
What this settles for the pages above. It is a vision-only, image-only result: there is no audio…
- GLM (Z.AI)×3
Gemma 4 — the sibling open-weight family with the opposite strategy (small + efficient vs large +…
- Claude Fable 5×2
Everything above is vendor-claim. One external number now exists: DeepMind's Gemma 4 report…
- Compute-Controlled Benchmarking×2
Gemma 4 — the worked example: a thinking-mode model benchmarked against a non-thinking predecessor…
- Jagged Intelligence (Ghosts, Not Animals)×2
Karpathy's examples are jaggedness across tasks at fixed model. Gemma 4 supplies a measured…
- Kimi (Moonshot AI)×2
Kimi K2.6 · 1T total / 32B active. Arena Text rank 34, Elo 1460, in Gemma 4's June 2026 table — 48…
- Native Multimodal Modeling: Fusion Depth and I/O Duality×2
Gemma 4 — the survey lists 31B and E4B under vision-encoder-based fusion; the encoder-free 12B is…
- Open Questions Backlog×2
Gemma 4: Is the encoder-free 12B's dense-text degradation an artifact of the 35M projection doing…
- Open-Weight Elicitation Irreversibility×2
> Status: wiki synthesis, not a source claim. No source argues this. Brown makes the budget…
- The Open-Weight Frontier Gap×2
Table 4 of the Gemma 4 report is a snapshot of the open-weight landscape as of June 19, 2026,…
- Responsible Scaling Policy Evaluations×2
Gemma 4 (DeepMind, July 2026, Apache 2.0) makes the shape visible. It ships a thinking mode; its…
- The Bitter Lesson×2
Gemma 4 — removes encoders and adds inference-path structure in the same release, drawing the line…
- Autonomous Intrusion
The negative findings are load-bearing and worth stating as claims rather than facts: Hugging Face…
- Google AI & Economy ATLAS
Gemma 4 · Google Deepmind — the sibling Google artifacts in this wiki
- Interaction / Background Model Split
Training makes the token learnable rather than rule-based: the SFT mixture includes ~8.5k hours of…
- Large-Scale Test-Time Compute
Brown's thesis is stated in the numerator — spend more, get more. It has a corollary he doesn't…
- Entities — People, Orgs, Tools & Projects
Gemma 4 — Google DeepMind's July 2026 open-weight multimodal family (Apache 2.0): 2.3B–31B dense…
- Single-Rollout Optimization
The Bitter Lesson — SAO both removes structure (drops the group baseline, drops π_θ_old) and adds a…
- TML-Interaction-Small
Gemma 4 — arrives at the same encoder-free design two months later, from memory constraints rather…
Related articles
- The Open-Weight Frontier Gap
Arena Text, June 2026: the top closed model leads the best open model by 33 Elo and the best *dense* open model by 57;…
- Compute-Controlled Benchmarking
Noam Brown's critique: the single-number benchmark grid is broken because it ignores test-time compute — plot performan…
- Inference Efficiency as Capability
If capability is a function of inference budget, then cutting the cost of a token is capability work: Gemma 4's five le…
- Kimi (Moonshot AI)
Moonshot AI's open-weight Kimi line — K2.5/K2.6 as 1T-class MoEs already circulating in this corpus (Inkling's post-tra…
- Open-Weight Elicitation Irreversibility
A wiki-drawn synthesis of Brown and Gemma 4: if dangerous capability scales with inference budget, then an open-weight…
