資料來源#
摘要#
互動模型(interaction model) 是一種 原生 處理互動的模型——持續接收音訊、視訊與文字,並即時思考、回應與行動——而不是透過外部鷹架(VAD、輪次偵測、對話管理 harness)來模擬即時行為。由 Thinking Machines Lab 於 2026 年 5 月以研究預覽形式發布,首個模型名為 TML-Interaction-Small。
核心賭注:互動性應當與智能一同擴展。 如果互動是模型的一部分,那麼擴展模型就會讓它同時變得更聰明 且 成為更好的協作者。如果互動寄生在手工打造的 harness 裡,The Bitter Lesson 告訴我們:這套 harness 會被通用能力的成長所超越。
The thesis in one line#
「要讓互動性隨智能擴展,它必須成為模型本身的一部分。」
這是把 harness 收縮的論點(參見 Harness Shrinkage as Models Improve)套用到 互動層:VAD、輪次邊界預測、對話狀態機——這些全都「明顯不如模型本身聰明」——都應該消融進模型行為之中。一旦如此,那些 harness 無法 支援的能力(主動插話、邊聽邊說、對視覺線索做出反應)就會變成模型行為的特例,並隨規模一同進步。
What it unlocks (capabilities, not harness features)#
- 無縫的對話管理——模型隱性地追蹤說話者是在思考、讓出發言權、自我修正,還是邀請回應。不需要獨立的對話管理元件。
- 語音與視覺插話——模型會在情境需要時介入(「當我說錯話時打斷我」、「當我寫出 bug 時告訴我」),而不只在輪次結束時。
- 同時發聲——使用者與模型同時說話(即時翻譯)。
- 時間感知——對流逝時間有直接的感受(「我跑一英里花了多久?」)。
- 同時的工具呼叫/搜尋/生成式 UI——在聆聽與說話的同時,模型並行搜尋、瀏覽、生成 UI,並把結果重新編織回對話中。
關於這些如何被展示與衡量,參見 Full-Duplex Interaction 與 Interactivity Benchmarks。
Why turn-based interfaces are the bottleneck#
今天的模型「以單一執行緒體驗現實」:它們盲目地等待,直到使用者打完字/說完話;接著盲目地生成,直到完成或被打斷。對協作而言,這是一條狹窄的通道——它限制了一個人的知識、意圖與判斷有多少能傳達給模型,也限制了模型的工作有多少是可被看見的。原文的類比是:用電子郵件而非當面去化解一場關鍵的分歧。完整論述見 Turn-Based Interface Bottleneck。
Architecture (three load-bearing ideas)#
- Time-Aligned Micro-Turns——輸入與輸出都是連續的串流,以 200ms 為單位的區塊來處理/生成;沒有人為的輪次邊界。沉默、重疊與打斷都保留在上下文中。
- Interaction / Background Model Split——一個具時間感知的互動模型維持即時的臨場感;一個非同步的背景模型負責持續的推理、工具使用與更長時程的工作。互動模型以一份 豐富的上下文封包(完整對話,而非獨立的查詢)來委派任務,並在符合使用者當下動作的時機把結果交錯地織回來。淨效果是:「以非思考型模型的回應延遲,達成推理型模型的規劃、工具使用與 agentic 工作流程。」
- Encoder-Free Early Fusion——以最小的前處理取代龐大的獨立編碼器/解碼器:音訊以 dMel + 輕量嵌入的形式輸入;影像透過 hMLP 轉為 40×40 的 patch;音訊則經由 flow head 輸出。所有元件都與 transformer 一同從零協同訓練。
再加上工程面:用於低開銷、高頻率小型 prefill/decode 的串流工作階段(已上游合併進 SGLang);針對延遲調校的 MoE kernel(以 gather+gemv 取代 grouped gemm);透過 batch-invariant kernel(<5% 開銷)達成訓練器與取樣器之間的位元級對齊,以確保穩定性與可除錯性。
Safety angle#
即時互動對安全性帶來的壓力,與以輪次為基礎的交流不同。TML 的工作聚焦於兩條軸線:
- 適配模態的拒絕——以 TTS 生成的拒絕/過度拒絕訓練資料,讓口說的拒絕既口語化、卻同樣堅定。
- 長時程的穩健性——自動化的紅隊 harness 生成多輪拒絕資料,維持與文字模型拒絕行為的一致性。
Limitations (per the post)#
- 長工作階段——連續的 A/V 會快速累積上下文;串流工作階段的設計能妥善處理短到中等長度的情況,但非常長的工作階段需要謹慎的上下文管理(與 Context Window Smart Zone 相呼應)。
- 算力與連線——低延遲的 A/V 串流需要可靠的連線;少了它就會嚴重劣化。
- 規模——
TML-Interaction-Small是 276B MoE / 12B active;更大的預訓練模型在今天這套機制下太慢、無法服務;更大的模型則承諾「今年稍晚」推出。 - 背景 agent——agentic 智能被承認是必要的,但相對於即時方面的工作,仍探索不足。
相關連結#
- Software 3.0——Karpathy 把互動模型視為邁向 3.0 神經電腦的一步
- The Bitter Lesson——整套方法所依憑的原則:擴展後的通用方法勝過手工打造的結構
- Harness Shrinkage as Models Improve——同樣的手法,套用在互動 harness 上(VAD/輪次偵測消融進模型)
- Turn-Based Interface Bottleneck——正在被解決的問題
- Time-Aligned Micro-Turns、Interaction / Background Model Split、Encoder-Free Early Fusion——三根架構支柱
- Full-Duplex Interaction——這所促成的各種互動模式
- Interactivity Benchmarks——如何把智能+互動性合在一起衡量
- Thinking Machines Lab、TML-Interaction-Small——誰打造了它、模型是什麼
- Context Window Smart Zone——長工作階段的侷限與 smart-zone 問題相呼應
- AI Employee Framing / Human-AI Accountability Redesign——兩者都反對純粹為自主性最佳化;互動模型是讓人類留在迴圈中的 介面側 解答
- Design Concept Grilling——以協作、即時的迭代取代「給規格然後走人」;互動模型正是能讓這種拷問式協作感覺起來原生的底層基礎
- Agent Harness Engineering——harness 與模型之間的分工問題,在這裡就互動層而言被明確地導向模型那一側
- Claude Opus 4.7——
xhigh努力等級作為一項基線設定出現(GPT-realtime-2.0 minimal/xhigh) - HTML as the New Markdown——對「更好的人類與 AI 協作」的一個姊妹解答,但採用相反的機制:TML 把 即時 介面消融進模型,而 Thariq Shihipar 則豐富 非同步 的產物;兩者都讓人類留在迴圈中
待解決的問題#
- 互動/背景的拆分是否能一般化,還是只是個過渡性產物,直到單一模型同時夠快又夠深為止?
- 「互動性隨智能擴展」只是個斷言;2026 年稍晚更大模型的發布,就是對它的檢驗。
- 已宣布針對互動性基準的研究經費——什麼會成為視訊主動性領域中 FD-bench 的對應物?
資料來源#
Cited by 27
- The Future of Agent Interfaces×5
Human collaboration · Interaction Models / Full Duplex Interaction · Human senses, speech, screen,…
- Opinions on Using AI Tools & the Future of the Software Engineering Role×3
Interface, not just code. Interaction Models / Turn Based Interface Bottleneck: today's turn-taking…
- Open Questions Backlog×3
Interaction Models: "Interactivity scales with intelligence" is asserted; the larger-model release…
- Thinking Machines Lab×3
Interaction Models (May 2026 research preview) — models that natively take in audio/video/text and…
- Build for the Next Model×2
The over-shoot he warns about — "too AGI-pilled for the moment." Ambrosino names the failure mode…
- Encoder-Free Early Fusion×2
A multimodal-design choice in Interaction Models: instead of routing audio and video through large,…
- Full-Duplex Interaction×2
"Full-duplex" = the model perceives and responds at the same time, in a constant two-way exchange —…
- HTML as the New Markdown×2
At first glance this contradicts the wiki's running Harness Shrinkage As Models Improve thesis (Cat…
- Interaction / Background Model Split×2
Interaction Models are architected as two cooperating models:
- The Bitter Lesson×2
Interaction Models — TML cites "the bitter lesson" directly: hand-crafted interactivity systems…
- Time-Aligned Micro-Turns×2
The core architectural move in Interaction Models: instead of consuming a complete user turn and…
- TML-Interaction-Small×2
Thinking Machines Lab's first interaction model — released as a research preview, May 2026. Pitched…
- Turn-Based Interface Bottleneck×2
Thinking Machines Lab's framing of why current AI interfaces limit collaboration: the turn-based…
- Agent Harness Engineering
Interaction Models — resolves the harness-vs-model question firmly toward the model for the…
- AI Employee Framing
Collaboration substrate: Interaction Models — real-time multimodal interaction as the interface…
- Anthropic
Thinking Machines Lab — peer lab; convergent harness-dissolves-into-model thesis (Interaction…
- Configurable Human Participation
Interaction Models — the channel-and-timing taxonomy (clarify-before-commit, feedback-after, mixed…
- Context Window Smart Zone
Interaction Models — continuous audio/video at 200ms granularity accumulates context fast; TML…
- Design Concept Grilling
Interaction Models — grilling is collaborative real-time iteration; turn-based interfaces are…
- Harness Shrinkage as Models Improve
Interaction Models — the same move on the interaction axis: VAD / turn-detection /…
- Human-AI Accountability Redesign
Interface-side complement: Interaction Models — org redesign keeps humans accountable; interaction…
- Inkling
Inkling-Small is a 276B MoE with 12B active — exactly the shape of Tml Interaction Small, TML's May…
- Interactivity Benchmarks
FD-bench, Audio MultiChallenge + new TimeSpeak/CueSpeak (proactive audio) and RepCount-A/ProactiveVideoQA/Charades (vis…
- Live-Path Minimalism
Interaction Models — the research framing whose serving problem this is
- Interaction & Multimodal
Interaction Models — Thinking Machines Lab (May 2026): models that handle audio/video/text…
- Software 3.0
Interaction Models — Thinking Machines' "diffusion-rendered UI / neural computer" direction is a…
- Thariq Shihipar
Interaction Models — both target better human–AI collaboration; his HTML artifacts enrich the…
Related articles
- Turn-Based Interface Bottleneck
Why current AI interfaces limit collaboration: single-thread turn-taking is a bandwidth bottleneck; humans pushed out b…
- Harness Shrinkage as Models Improve
Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…
- Interaction / Background Model Split
Dual-model architecture: time-aware interaction model stays present; async background model handles deep reasoning/tool…
- The Bitter Lesson
Sutton 2019: scaled general methods beat hand-engineered structure; recurring justification across the wiki for dissolv…
- Agent Harness Engineering
Patterns for scaffolding long-running LLM agents: environment design, progressive context disclosure, mechanical archit…
