資料來源#
這是什麼#
Thinking Machines Lab 的首個**互動模型——於 2026 年 5 月以研究預覽版形式發布。其定位是「首個同時具備強大智慧/指令遵循能力以及**互動性的模型」。
- **架構:**276B 參數 MoE,12B active。從零開始以互動模型的身分訓練(而非在回合制模型上外加互動性)。
- **模態:**連續音訊+視訊+文字輸入;文字+音訊輸出。Encoder-Free Early Fusion(dMel 音訊嵌入、用於影格的 40×40-patch hMLP、用於音訊輸出的 flow head),單一共享 transformer,所有元件皆從零開始共同訓練。
- 互動機制:Time-Aligned Micro-Turns——200ms 交錯輸入/輸出區塊,不設回合邊界。
- **推理:**將深度推理/工具使用/長時間跨度工作委派給非同步背景模型——參見 Interaction / Background Model Split。即使沒有背景 agent,在智慧基準測試上的表現也具競爭力。
主要數據(2026 年 5 月)#
- 輪替延遲:0.40s(FD-bench v1,音訊)——在所有比較的模型中最佳。
- FD-bench v1.5 平均分:77.8,相較之下,包括 thinking-high models 在內的基準模型約為 39–54。
- FD-bench v3(音訊+工具):82.8% 回應品質/68.0% Pass@1(搭配背景 agent)。
- Audio MultiChallenge APR:43.4%——擊敗每個 non-thinking baseline;只有 GPT-realtime-2.0 xhigh(48.5%)更高。
- 比較的基準模型:GPT-realtime-2.0(minimal/xhigh)、GPT-realtime-1.5、Gemini-3.1-flash-live-preview(minimal/high)、Qwen 3.5 Omni-plus-realtime。完整表格見互動性基準測試。
限制(已承認)#
- 長時間連續的 A/V 工作階段會快速累積上下文——謹慎的上下文管理仍是未解決的問題(呼應上下文視窗智慧區)。
- 需要可靠的低延遲連線;缺乏連線時效能會嚴重下降。
- 之所以稱為「Small」,是因為更大型的預訓練模型目前服務速度太慢,無法適用於此運作模式——承諾在 2026 年稍晚推出更大型模型。
可用性#
有限的研究預覽版將於「未來幾個月內」推出,更廣泛的發布則在「今年稍晚」。歡迎透過 interaction@thinkingmachines.ai 提供意見;研究補助申請已開放。
相關連結#
- 互動模型——模型類別
- Thinking Machines Lab——建造者
- Time-Aligned Micro-Turns/Encoder-Free Early Fusion/Interaction / Background Model Split——其三大架構支柱
- Full-Duplex Interaction——它所展示的互動模式
- 互動性基準測試——完整基準測試表格,以及它擊敗的基準模型
- Claude Opus 4.7——兩者同屬同時代模型(2026 年年中前沿模型);4.7 的
xhigheffort tier 對應 GPT-realtime 的 minimal/xhigh,本文將其作為基準設定 - 上下文視窗智慧區——長工作階段的限制
- Gemma 4——兩個月後採用相同的無編碼器設計,但其動機來自記憶體限制而非延遲;它是 TML 架構押注最接近的獨立複製
資料來源#
Cited by 13
- Inkling×4
Inkling-Small is a 276B MoE with 12B active — exactly the shape of Tml Interaction Small, TML's May…
- Thinking Machines Lab×3
Interaction Models (May 2026 research preview) — models that natively take in audio/video/text and…
- Interaction / Background Model Split×2
At the split's introduction the background model was an unnamed capability. Inkling fills the slot:…
- Interaction Models×2
An interaction model is a model that handles interaction natively — continuously taking in audio,…
- Interactivity Benchmarks×2
The evaluation surface Thinking Machines Lab uses to argue Tml Interaction Small is "the first…
- Claude Opus 4.7
Tml Interaction Small — era-mate (mid-2026 frontier from a different lab); 4.7's xhigh effort tier…
- Encoder-Free Early Fusion
Tml Interaction Small — the model that implements this design (dMel audio, 40×40 hMLP patches, flow…
- Full-Duplex Interaction
Tml Interaction Small — the model that demonstrates these interaction modes
- GPT-Live
Tml Interaction Small — the research-preview sibling: same architectural conclusions from Thinking…
- Kimi (Moonshot AI)
K3 ships a 401M MoonViT-V2 vision encoder at 2.8T scale. Encoder Free Early Fusion documents three…
- Entities — People, Orgs, Tools & Projects
Tml Interaction Small — TML's first interaction model: 276B MoE / 12B active, audio+video+text in /…
- Open Questions Backlog
Inkling ×3 (oldest 21d) — Inkling-Small's 276B/12B dimensions match Tml Interaction Small exactly.…
- Time-Aligned Micro-Turns
Tml Interaction Small — the model built on this mechanism (200ms interleaved input/output chunks)
Related articles
- Interaction Models
Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…
- Encoder-Free Early Fusion
Multimodal design with minimal pre-processing instead of large standalone encoders: TML co-trains dMel audio + 40×40-pa…
- Interaction / Background Model Split
Dual-model architecture: time-aware interaction model stays present; async background model handles deep reasoning/tool…
- Full-Duplex Interaction
Perceive-and-respond simultaneously across modalities; proactive interjection, visual-cue reactions, simultaneous speec…
- Inkling
Thinking Machines Lab's first from-scratch open-weights family (July 2026): a 975B/41B-active multimodal MoE with 1M co…
