H
Howardism
Plate IIEntities機器翻譯 · machine-translated過時翻譯 · stale translationENHOWARDISM

TML-Interaction-Small

PublishedMay 13, 2026FiledEntityDomainEntitiesTagsType/entityLLM ModelReading3 minSourceAI-synthesised

TML 的首個互動模型:276B MoE/12B active, 音訊+視訊+文字輸入;文字+音訊輸出;200ms 微回合;非同步背景 agent;所有模型中最佳的輪替延遲;2026 年 5 月研究預覽版

TML-Interaction-Small 插圖

資料來源#

這是什麼#

Thinking Machines Lab 的首個**互動模型——於 2026 年 5 月以研究預覽版形式發布。其定位是「首個同時具備強大智慧/指令遵循能力以及**互動性的模型」。

  • **架構:**276B 參數 MoE,12B active。從零開始以互動模型的身分訓練(而非在回合制模型上外加互動性)。
  • **模態:**連續音訊+視訊+文字輸入;文字+音訊輸出。Encoder-Free Early Fusion(dMel 音訊嵌入、用於影格的 40×40-patch hMLP、用於音訊輸出的 flow head),單一共享 transformer,所有元件皆從零開始共同訓練。
  • 互動機制:Time-Aligned Micro-Turns——200ms 交錯輸入/輸出區塊,不設回合邊界。
  • **推理:**將深度推理/工具使用/長時間跨度工作委派給非同步背景模型——參見 Interaction / Background Model Split。即使沒有背景 agent,在智慧基準測試上的表現也具競爭力。

主要數據(2026 年 5 月)#

  • 輪替延遲:0.40s(FD-bench v1,音訊)——在所有比較的模型中最佳。
  • FD-bench v1.5 平均分:77.8,相較之下,包括 thinking-high models 在內的基準模型約為 39–54。
  • FD-bench v3(音訊+工具):82.8% 回應品質/68.0% Pass@1(搭配背景 agent)。
  • Audio MultiChallenge APR:43.4%——擊敗每個 non-thinking baseline;只有 GPT-realtime-2.0 xhigh(48.5%)更高。
  • 比較的基準模型:GPT-realtime-2.0(minimal/xhigh)、GPT-realtime-1.5、Gemini-3.1-flash-live-preview(minimal/high)、Qwen 3.5 Omni-plus-realtime。完整表格見互動性基準測試

限制(已承認)#

  • 長時間連續的 A/V 工作階段會快速累積上下文——謹慎的上下文管理仍是未解決的問題(呼應上下文視窗智慧區)。
  • 需要可靠的低延遲連線;缺乏連線時效能會嚴重下降。
  • 之所以稱為「Small」,是因為更大型的預訓練模型目前服務速度太慢,無法適用於此運作模式——承諾在 2026 年稍晚推出更大型模型。

可用性#

有限的研究預覽版將於「未來幾個月內」推出,更廣泛的發布則在「今年稍晚」。歡迎透過 interaction@thinkingmachines.ai 提供意見;研究補助申請已開放。

相關連結#

資料來源#

§ end
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Cited by 13
  • Inkling×4

    Inkling-Small is a 276B MoE with 12B active — exactly the shape of Tml Interaction Small, TML's May…

  • Thinking Machines Lab×3

    Interaction Models (May 2026 research preview) — models that natively take in audio/video/text and…

  • Interaction / Background Model Split×2

    At the split's introduction the background model was an unnamed capability. Inkling fills the slot:…

  • Interaction Models×2

    An interaction model is a model that handles interaction natively — continuously taking in audio,…

  • Interactivity Benchmarks×2

    The evaluation surface Thinking Machines Lab uses to argue Tml Interaction Small is "the first…

  • Claude Opus 4.7

    Tml Interaction Small — era-mate (mid-2026 frontier from a different lab); 4.7's xhigh effort tier…

  • Encoder-Free Early Fusion

    Tml Interaction Small — the model that implements this design (dMel audio, 40×40 hMLP patches, flow…

  • Full-Duplex Interaction

    Tml Interaction Small — the model that demonstrates these interaction modes

  • GPT-Live

    Tml Interaction Small — the research-preview sibling: same architectural conclusions from Thinking…

  • Kimi (Moonshot AI)

    K3 ships a 401M MoonViT-V2 vision encoder at 2.8T scale. Encoder Free Early Fusion documents three…

  • Entities — People, Orgs, Tools & Projects

    Tml Interaction Small — TML's first interaction model: 276B MoE / 12B active, audio+video+text in /…

  • Open Questions Backlog

    Inkling ×3 (oldest 21d) — Inkling-Small's 276B/12B dimensions match Tml Interaction Small exactly.…

  • Time-Aligned Micro-Turns

    Tml Interaction Small — the model built on this mechanism (200ms interleaved input/output chunks)

Related articles
  • Interaction Models

    Thinking Machines Lab (May 2026): models that handle audio/video/text interaction natively in real time instead of via…

  • Encoder-Free Early Fusion

    Multimodal design with minimal pre-processing instead of large standalone encoders: TML co-trains dMel audio + 40×40-pa…

  • Interaction / Background Model Split

    Dual-model architecture: time-aware interaction model stays present; async background model handles deep reasoning/tool…

  • Full-Duplex Interaction

    Perceive-and-respond simultaneously across modalities; proactive interjection, visual-cue reactions, simultaneous speec…

  • Inkling

    Thinking Machines Lab's first from-scratch open-weights family (July 2026): a 975B/41B-active multimodal MoE with 1M co…