資料來源#
這是什麼#
Thinking Machines Lab 首次從零開始訓練並發布的模型(2026 年 7 月),完整權重可於 Hugging Face 取得(原始 checkpoint 及適用於 Blackwell 的 NVFP4 版本)。這是一個 Mixture-of-Experts transformer,總計 975B/啟用 41B,context 最長可達 1M tokens,以 45T tokens 的文字、影像、音訊與影片資料預訓練。另同步預覽 Inkling-Small——總計 276B/啟用 12B,採用相同配方,待測試完成後承諾釋出權重。
它的定位明確而少見:「Inkling 並非目前整體最強的模型,無論開放或封閉皆然。」 它主打成為最適合客製化的基礎模型——具多模態能力、效率高,並可在 Tinker 上微調;Tinker 是 TML 託管的微調平台(提供 64K/256K context 選項)。這個模型的存在是為了供應平台:公告的核心展示是 Inkling 在 Tinker 上微調自身(撰寫、執行並評估自己的不含特定字母訓練工作),而這次釋出也附有 cookbook 配方、Playground 聊天介面,以及用於透過工具呼叫和多模態輸入進行取樣/後訓練的 tml-renderer 函式庫。這是第三種開放權重策略,有別於《開放權重前沿差距》中的兩極:不是以稀疏化逼近前沿(GLM、Kimi、DeepSeek),也不是以效率適配裝置(Gemma 4),而是以適應性服務微調者。
架構#
與常見配方不同之處,各自都以效率或長 context 效能作為理由:
- MoE 大致沿用 DeepSeek-V3:每層有 256 個路由專家和 2 個共享專家,每個 token 啟用 6 個路由專家,使用 sigmoid router 與無輔助損失的負載平衡;路由專家與共享專家的分數會共同正規化。
- Attention:滑動視窗層與全域層以 5:1 交錯排列,搭配 8 個 KV heads——和 Gemma 4 為節省 KV-cache 而選用的比例相同,現在有第二間實驗室採用。使用相對位置嵌入而非 RoPE;TML 表示這在長 context 下表現更好、外插能力也更佳——這與業界共識不同,值得觀察能否重現。
- 短卷積:置於 key/value 投影之後,以及 attention/MLP 殘差分支輸出上。
- 無 encoder 多模態:音訊以 dMel spectrograms 輸入,影像則切成 40×40 像素區塊,經由四層 hMLP 處理;兩者都透過輕量 embedding 層進入共享 transformer——這是無 Encoder 早期融合的第三個案例,也是 TML 自家的第二個案例(繼 TML-Interaction-Small之後),這次達到 975B 開放權重規模。
訓練#
在 NVIDIA GB300 NVL72 系統上預訓練,採用混合式 optimizer——大型矩陣用 Muon,其餘用 Adam——並將 weight decay 與學習率平方耦合(源自 TML 的 modular-manifolds 研究);他們表示這讓不同訓練時程下的權重範數保持穩定。
後訓練先以開放權重模型(包括 Kimi K2.5)生成的合成資料進行 SFT,作為引導起始——這個從零開始預訓練的模型,其後訓練種子語料部分來自競爭者的開放模型;據稱這個引導階段只占少部分算力。大部分算力用於大規模非同步 RL:在兩段長時間連續訓練中完成超過 3,000 萬次 rollout,留出的推理綜合指標(AIME、HLE、GPQA)全程呈對數線性提升(廠商圖表)。
兩項比基準測試更值得關注的 RL 發現:
- 可控的思考力度是訓練出來的,不是靠腳手架實現:每筆樣本透過 system message 指定力度等級,並調整每 token 成本,藉此教出連續的力度調節(0.2–0.99),可在 harness 內設定。
- 湧現的 CoT 壓縮:RL 訓練期間,思路鏈變得電報式精簡——省略冠詞與連接詞(「We need to understand」→「We need determine」)——訓練中並未對此給予獎勵;只有效率壓力便促成壓縮。Cognition 回報,SWE-1.7 訓練期間也出現相同現象。這項現象記錄於思路鏈可監測性,其監測性影響也在該文討論。
表現定位(廠商回報,effort=0.99)#
在開放 MoE 巨型模型中能力居中、廣度領先,安全性與 token 效率也出色:
- 在高難度推理與程式編寫上落後 GLM 5.2/Kimi K2.6:HLE 純文字 29.7%(GLM 5.2:40.1)、Terminal Bench 2.1 63.8%(GLM 5.2:82.7)、SWE-bench Pro 54.3%(GLM 5.2:62.1)。
- 在以下項目領先受比較的開放模型:MCP Atlas(74.1%)、IFBench(79.8%——高於 TML 自家表格中的所有封閉模型,包括 Claude Fable 5 的 63.5)、SimpleQA Verified(43.9%)、音訊(AudioMC 56.6%,開放全能型專門模型為 24–38;VoiceBench 91.4%)。
- 力度曲線:Terminal Bench 成績追平 Nemotron 3 Ultra,但 token 數約只有三分之一。TML 公布完整的力度/效能掃描結果,並與競爭者的預設運作點比較——這是最接近算力控制基準測試所要求「以曲線而非網格呈現」的廠商發布,雖然只有自家模型有完整曲線。
- 安全性:FORTRESS Adversarial 78.0%——在受比較的開放權重模型中內建防護最強——對良性相似案例的辨識率為 95.9%;StrongREJECT 98.6%。外部測試者檢驗了 CBRN/網路安全/失控,以及諂媚/脆弱使用者/操弄等面向。和 Gemma 4 未列入表格的安全性說明不同,這裡公布了數字——開放權重誘發不可逆性說明單點評估開放權重模型能、不能證明什麼。
- 校準:ForecastBench Brier Index 61.1(未搜尋)——與 Gemini 3.1 Pro 相當,高於 GPT-5.5(59.1),也遠高於 TML 表格中的 Claude Opus 4.8(54.6)。背後的訓練配方見訓練式校準。
- 基準測試方法說明:評估時 effort 設為 0.99、temperature 設為 1.0、trajectory 上限為 256K tokens;有外部回報(Artificial Analysis)數據時便採用;經網路搜尋發現受污染的 Terminal Bench rollout 一律計 0 分——這是自願遵循基準污染與去污染所提倡紀律的小型案例。
以上全部屬於 vendor-claim:數據來自 TML 自家表格,其中部分使用預發布 checkpoint(ForecastBench 結果於 2026 年 6 月 30 日至 7 月 13 日以不同於最終釋出版本的 checkpoint 執行)。
Inkling-Small 與互動模型的關聯#
Inkling-Small 是一個啟用 12B 的 276B MoE——規模與 TML 於 2026 年 5 月推出的互動模型 TML-Interaction-Small完全相同。公告指出,Inkling 的設計目標是作為互動模型系統中的背景推理模型(互動/背景模型拆分)。合併來看,TML 雙模型架構中即時快速的一半(TML-Interaction-Small)與深入非同步的一半(Inkling)現在都已公開,而且小型版的規格與互動模型相同。Inkling-Small 在幾項基準上追平或超過完整版 Inkling(使用工具的 HLE 為 46.6 對 46.0、IFBench 為 83.4 對 79.8、GPQA 為 88.3 對 87.2)——TML 將此歸因於大型訓練完成後改進的預訓練配方;但在規模有助於儲存知識的項目上則落後(SimpleQA 為 20.9 對 43.9)。
生態系#
服務堆疊的首日合作夥伴包括:Together AI、Fireworks、Modal、Databricks、Baseten 提供 API;SGLang 和 Miles(RadixArk)、vLLM(Inferact)、TokenSpeed(Lightseek)、llama.cpp(Unsloth)支援推論/RL;並整合 Hugging Face transformers。Tinker 推出 Inkling 時提供五折優惠。
相關連結#
- Thinking Machines Lab——開發實驗室;這是它從研究預覽轉向正式推出開放權重模型
- TML-Interaction-Small——規格與 Inkling-Small 同為 276B/12B 的同系列模型;是這對模型中的互動端
- 互動/背景模型拆分——Inkling 是拆分架構中明確命名的背景推理端
- 無 Encoder 早期融合——此設計的第三個案例,也是首個約 1T 規模的開放權重案例
- 《開放權重前沿差距》——在前沿開放 MoE 榜單中加入西方廠商,競爭重點是客製化而非峰值能力
- 訓練式校準——認知訓練配方(proper scoring rules、納入棄答考量的獎勵、雙評分器)
- 思路鏈可監測性——推理軌跡在 RL 訓練期間湧現的壓縮現象
- 算力控制基準測試——公布自家的力度/效能曲線;競爭者仍只以單點呈現
- 開放權重誘發不可逆性——安全數據有表列,且提供託管微調途徑,讓開放權重安全問題更複雜,但尚未解決
- Gemma 4——2026 年另一個西方開放權重模型,採取相反策略,attention 比例同樣為 5:1
- Kimi (Moonshot AI)——其 K2.5 為 Inkling 的後訓練提供種子資料;而同月數日後推出的 K3(2.8T/104B)取得 Inkling 選擇不追逐的前沿開放規模紀錄,並保留 Inkling 捨棄的視覺 encoder
- 《開放權重作為競爭策略》——此文引用 Inkling 作為證據,未深入介紹。 Andrew Ng 將它(與 NVIDIA 的 Nemotron)稱為美國能在開放模型領域競爭的證明——「表現非常亮眼」——並主張開放權重是提升國家競爭力的工具。Inkling 強調可微調,也支持他的次要主張:所謂「最佳可建構基礎」談的是誰能掌握應用層,而非排行榜名次。該文也提出蒸餾公平性論點,而此模型正好是個棘手案例:上述 K2.5 引導階段屬於跨實驗室蒸餾,方向與東西方常見方向相反,且由蒸餾方公開說明
尚待釐清的問題#
- Inkling-Small 的 276B/12B 規格與 TML-Interaction-Small完全相同。互動模型是從 Inkling 系列微調而來(或反之)嗎?TML 會把拆分的兩端整合成同一個模型家族嗎?
- 後訓練以 Kimi K2.5 合成資料作為引導。使用競爭者資料作為引導,會留下可量測的痕跡(風格、拒絕模式、tokenizer 慣用語回聲),並在 3,000 萬次 RL rollout 後留存嗎?
- TML 聲稱相對位置嵌入在長 context 外插方面勝過 RoPE,與目前業界共識相左。這項主張能在 TML 以外的環境、1M context 下重現嗎?
資料來源#
- Inkling: Our Open-Weights Model — TML 公告(2026-07,
vendor-claim):架構、訓練、基準測試表、力度掃描、安全數據、生態系
Cited by 18
- Kimi (Moonshot AI)×5
Kimi K2.6 · 1T total / 32B active. Arena Text rank 34, Elo 1460, in Gemma 4's June 2026 table — 48…
- Compute-Controlled Benchmarking×4
The defection is half because it is asymmetric: competing models appear at their default operating…
- Chain-of-Thought Monitorability×3
The Korbak worry assumed the pressure comes from training on the trace. Inkling's training report…
- Encoder-Free Early Fusion×3
Inkling (vendor-claim) carries the design from a 276B research preview and a 12B edge model to a…
- Thinking Machines Lab×3
Inkling — their first from-scratch, full-weights release (July 2026); the customization thesis made…
- Native Multimodal Modeling: Fusion Depth and I/O Duality×2
Nothing from the interaction-model line appears at all: no Tml Interaction Small, no Inkling, no…
- Open-Weight Elicitation Irreversibility×2
Inkling is the disclosure counter-example to premise 3 while leaving the argument untouched. Where…
- The Open-Weight Frontier Gap×2
A third strategy arrives: customize, don't compete (July 2026). Inkling — TML's 975B/41B-active…
- TML-Interaction-Small×2
Inkling-Small — the preview sibling of TML's open-weights release — is a 276B MoE with 12B active:…
- Trained Calibration×2
Calibration — expressing the right amount of confidence, including on unsettled questions — treated…
- Autonomous Intrusion
The negative findings are load-bearing and worth stating as claims rather than facts: Hugging Face…
- Cheating in Capability Evaluations
Opus 4.7: the trace is absent, and the cause is an efficiency feature. AISI's explanation is that…
- Interaction / Background Model Split
At the split's introduction the background model was an unnamed capability. Inkling fills the slot:…
- Interaction Models
The limit worth stating. The survey's census is restricted to open-source models and technical…
- Entities — People, Orgs, Tools & Projects
Inkling — Thinking Machines Lab's first from-scratch open-weights family (July 2026): a…
- Open Questions Backlog
Inkling ×3 (oldest 75d) — Inkling-Small's 276B/12B dimensions match Tml Interaction Small exactly.…
- Open Weights as Competitive Strategy
Ng points to two domestic open releases as evidence the US can compete on this axis without…
- Task Gaming
Inkling — measured on three environments here, including the largest clean human-in-the-loop drop…
Related articles
- Kimi (Moonshot AI)
Moonshot AI's open-weight Kimi line — K2.5/K2.6 as 1T-class MoEs already circulating in this corpus (Inkling's post-tra…
- Gemma 4
Google DeepMind's July 2026 open-weight multimodal family (Apache 2.0): 2.3B–31B dense plus a 26B/4B-active MoE, adding…
- The Open-Weight Frontier Gap
Arena Text, June 2026: the top closed model leads the best open model by 33 Elo and the best *dense* open model by 57;…
- Encoder-Free Early Fusion
Multimodal design with minimal pre-processing instead of large standalone encoders: TML co-trains dMel audio + 40×40-pa…
- Inference Efficiency as Capability
If capability is a function of inference budget, then cutting the cost of a token is capability work: Gemma 4's five le…
