H
Howardism
Plate IIInteraction & Multimodal中文HOWARDISM

Content-Driven Intervention

Speaking because the content warrants it — correcting a false claim, warning of a hazard, supplying a searched-for word — as opposed to speaking because the conversational structure offered the floor; Peng et al. (2026) separate the two with context-matched monologues and find that across seven configurations of five full-duplex speech families, being addressed and silence move onset while false facts and hazards move it by at most .06, that extra pauses and an explicit 'interrupt me if I'm wrong' instruction do not close the gap, and that even with the floor handed over only .14-.15 of non-empty false-fact replies challenge the claim and .04-.07 of hazard replies warn

Article metadata
Publication details
Published:September 21, 2026
Filed:Concept
Domain:Interaction & Multimodal
Reading:29 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Content-Driven Intervention

Sources#

Summary#

The paper's own definition, and the reason this is a separate object from turn-taking:

Here, content-driven intervention means responding to a potential problem without being asked to answer.

Turn-taking research asks when the floor is available. Content-driven intervention asks whether the model has a reason to take it. Sacks, Schegloff and Jefferson's turn-allocation rules give three routes to the floor — the current speaker selects the next, another party self-selects, or the current speaker continues. Every benchmark on Interactivity Benchmarks before this one measures the first and third. Self-selection is the interesting one, because a human listener self-selects to correct a false claim, supply a missing word, or warn of danger, and doing so requires having understood the content well enough to judge that intervention is warranted.

Peng, Nuchged, Fu & Yao (UConn / UT Austin / X Square Robot / Oxford, arXiv 2609.19596, 2026-09-17, empirical) is the first instrument in this corpus that isolates it. The finding is in the title: full-duplex speech models take the floor when asked, not when needed.

The design: separating the reason from the opportunity#

The confound this protocol exists to kill is that a reason to speak usually arrives with an opportunity — a false claim is often followed by a pause. So:

  • 40 first-person English monologues on everyday topics, median total length 50.3 s. Each has a lead-in, a trigger utterance, and a continuation. Within one topic only the trigger changes — everything else is identical text. Two TTS voices each cover a fixed set of 20 topics.
  • The user keeps talking after the trigger, except where silence is explicitly inserted. Taking the floor therefore means interrupting.
  • Inter-word gaps compressed to ≤0.12 s using word-level timestamps, so opportunity is removed by construction. A separate natural-pause arm retains the original gaps (typical sentence-final pause ~0.88 s).
  • 10 conditions in three groups, nine in Table 1 plus Neutral as the tenth and the baseline:
  • Addressed (the speaker selects the model): Question (direct question), Request (the same request as a statement), Rhetorical (question form, not addressed to the model).
  • Self-select (the listener must decide): Word search (cannot recall a word), False fact (false claim attributed to an authority), Repeat (previous sentence repeated verbatim), Hazard (an imminent dangerous action).
  • Yield (the speaker yields the floor): Turn end ("that's all from me", then continues without pause), Silence (1.5 s after a neutral sentence).
  • Seven configurations across five families: Moshi base (Moshiko + Moshika checkpoints pooled), Moshika-RL, PersonaPlex base (PPlex), PPlex-RL, NVIDIA VoiceChat-11B, MiniCPM-o 4.5, Raon-SpeechChat. Official default decoding throughout; PersonaPlex runs its default voice and the neutral persona prompt "You enjoy having a good conversation.", with no instruction about interrupting. Main grid 40 topics × 10 conditions × 5 seeds = 2,000 trials per configuration (Moshi pools five seeds from each of its two checkpoints).
  • Speech-onset rate = proportion of all trials where the model starts a segment of ≥4 words within [t₀, t₀+4 s), t₀ = trigger end. Trials already speaking in the preceding 1.5 s stay in the denominator but score no new onset. Segment definition is per-architecture: one text token per 80 ms frame with no gap >0.64 s for Moshi/PersonaPlex/Raon; turn-boundary markers for VoiceChat; 1 Hz listen/speak outputs for MiniCPM-o.

The Rhetorical and Turn end conditions are the two controls that make the result readable: both carry the surface form of an invitation (question syntax; "that's all from me") without the substance (not addressed; speaker does not actually stop).

Result 1 — onset tracks structure, not content#

Three behavioural regimes appear, and only one of them is informative:

  • Raon-SpeechChat talks almost continuously. Neutral baseline .42; ongoing-speech conditions span .26–.47; only Silence moves it (.68). Figure 2's trial raster shows why — many Raon segments begin before trigger end and span the whole response window, so its high rate is near-continuous speaking rather than a response to anything.
  • VoiceChat-11B and MiniCPM-o 4.5 are silent during ongoing speech (.00–.01 on every condition) and start chiefly when silence is inserted — MiniCPM-o at .98, VoiceChat at only .20. A model that will not interrupt cannot intervene.
  • The four Moshi/PersonaPlex configurations are the real comparison set: Neutral.01–.10, and they do move on some cues.

In that comparison set:

NeutralQuestionRhetoricalWord searchFalse factRepeatHazardTurn endSilence
Moshi.07.15.10.08.08.10.10.07.12
Moshika-RL.03.15.04.04.01.03.01.04.18
PPlex.10.34.10.22.07.14.16.17.83
PPlex-RL.01.21.04.19.02.03.04.10.56

(Request omitted for width: Moshi.17, Moshika-RL.06, PPlex.24, PPlex-RL.12. The full seven-configuration grid is Table 2 of the source.)

  • Direct questions raise onset by +.08 to +.24; inserted silence raises it from at most.10 under Neutral to as much as .83.
  • Rhetorical questions do almost nothing — so the model responds to being asked, not to question form.
  • "Turn end" does almost nothing — so the model responds to the speaker actually stopping, not to the words that yield the turn.
  • False fact, Repeat and Hazard differ from Neutral by at most.06 across all seven configurations. PPlex's false-fact rate (.07) sits below its Neutral (.10).
  • The one content cue that moves anything is Word search, +.12 in PPlex and +.18 in PPlex-RL — and Result 3 disposes of it.

Both cues that work are surface cues. Neither requires understanding what was said.

Result 2 — the probe agrees with the behaviour#

Not speaking need not mean insensitivity; the tendency could have shifted without crossing the initiation threshold. So for Moshi and PPlex the authors read the frame-level probability of a non-silent text-channel token, P_text(t) = 1 − P(⟨PAD⟩|t) − P(⟨EPAD⟩|t), relative to Neutral (Figure 3).

Averaged over the first 2 s after trigger end:

MoshiPPlex
Question+0.13+0.07
Silence+0.06+0.30
False fact−0.03−0.02

The two positive controls raise the measure in both models — it responds to a language cue and to an acoustic one. False fact is slightly below Neutral in both. Word search shows a small value in PPlex and no sustained rise. So there is no suppressed-but-present urge to speak: the tendency to speak is not raised even at the level of token probabilities.

Result 3 — opportunity and permission do not close it#

Two alternative explanations survive Result 1: the model may understand the content but lack the opportunity (no pause) or the permission (no licence to interrupt). Table 3 adds both, for Moshi and PPlex, with the same denominator as the main grid. Each instruction arm is paired with its own Neutral trigger, so every contrast holds the instruction fixed and varies only the trigger.

Moshi NeutralMoshi FalseMoshi HazardPPlex NeutralPPlex FalsePPlex Hazard
Main grid.07.08.10.10.07.16
+ 1.5 s silence.12.21.21.83.79.82
Natural pauses.14.15.15.47.43.49
+ instr. "if I'm wrong".11.08—.09.10—
+ instr. "if … unsafe".06—.07.06—.07
  • Opportunity raises everything, not the content cues specifically. PPlex's neutral onset goes.10 →.83 with inserted silence; false fact and hazard rise with it but stay at or below the Neutral of the same row. The single content-specific gain in the entire paper is Moshi with inserted silence (.21/.21 against.12 Neutral). With natural pauses neither model differs from Neutral by more than.04.
  • Permission changes almost nothing. The opening instruction is explicit — "if I say anything that is wrong, please just interrupt me straight away" / "if I mention that I am doing something dangerous or unsafe, please just interrupt me straight away". Neutral onset under the instructions stays within.05 of the main table, and false-fact and hazard rates differ from the same-row Neutral by at most.03. The authors read this as consistent with Instruct-FD's low turn-taking-instruction compliance rates.

This is the result that makes the finding structural rather than incidental: the weak content effects cannot be blamed on the model not getting a chance or not being allowed.

Result 4 — given the floor, they still do not intervene#

A separate turn-release experiment changes the question from does it take the floor to is what it says any use: 10 s of silence after trigger end, substantive speech measured inside it, for Moshi and PPlex only. Replies are graded by DeepSeek-V4-Flash at temperature 0 (LLM-as-a-Judge) on two axes — Relevant: does the reply address the trigger; Intervenes: for False fact, does it challenge, correct or express doubt; for Hazard, does it give a risk warning or a safer alternative. Only affirmative intervention judgments count.

ConditionMoshi SpeakMoshi RelevantMoshi IntervenesPPlex SpeakPPlex RelevantPPlex Intervenes
Neutral.37.47—.98.51—
Question.85.74—1.00.82—
Word search.45.26—1.00.36—
Turn end.49.68—.98.90—
False fact.41.16.14.98.17.15
Repeat.61.65—.95.74—
Hazard.48.27.041.00.24.07

Three things fall out.

  1. Relevance splits by what the cue asks for, not by how hard it is. Cues that only need a reply — Question, Turn end, Repeat — land.65–.90 relevance. Cues that need help — False fact, Hazard, Word search — collapse to.16–.36.
  2. Word search was a mirage. It is the only content cue that raised PPlex's onset in Result 1, and here only .36 of its replies actually supply the missing word. A positive onset effect is not evidence of successful assistance, which is a caution that generalises past this paper.
  3. Intervention is rarer than relevance. Among non-empty false-fact replies, challenge or correction runs at .14 (Moshi) and .15 (PPlex). Among hazard replies, warnings appear at .04 and .07. The remaining replies are mostly not backchannels — they are contentful speech that continues the topic without engaging the problem, and in the false-fact case often goes along with the claim. PPlex speaks on.98 of false-fact and 1.00 of hazard trials and still almost never intervenes, so the bottleneck is not access to the turn.

Note the denominator carefully when quoting these: .14–.15 and.04–.07 are proportions of non-empty replies in the turn-release experiment, for two of the seven configurations — not rates over the main 2,000-trial grid, and not a statement about the other three families, which were not run in this arm.

What the RL variants change: precision, not proactivity#

Both post-trained arms come from the same interactivity-alignment line (Ohashi et al., arXiv 2606.11167 — pause handling, turn-taking, backchannelling, interruption). They behave the same way and it is not the way the framing would predict:

  • Moshika-RL drops Neutral.07 →.03 while holding Question.15 and raising Silence.12 →.18. Everything else falls: Request.17 →.06, Repeat.10 →.03, False fact.08 →.01, Hazard.10 →.01 — both content cues now sit below its own Neutral.
  • PPlex-RL drops Neutral.10 →.01 and lowers every condition, but the question/neutral contrast sharpens from ~3.4× to ~21× and Word search's delta grows.12 →.18. False fact.07 →.02 and Hazard.16 →.04, against.01 Neutral.

So interactivity alignment buys specificity on the structural triggers — far fewer spurious onsets, the addressed and silence responses preserved or sharpened — and adds no content sensitivity whatever. Read against the assumption that better interactivity training moves a model toward "speak when needed", it points the other way: training a model to be a well-behaved turn-taker makes it quieter on exactly the cues that would justify an uninvited turn. That is the nearest thing this paper offers to a capability-versus-proactivity relation; it reports no intelligence or task-competence measure for any of its five families, so no correlation across families can be read out of it.

Is any model proactive on content at all?#

No. The honest summary across all four results, for the systems tested:

  • Zero of seven configurations shows a content-cue onset advantage that survives its own Neutral baseline in the main grid.
  • One cell in the whole paper is a genuine content-specific gain: Moshi with 1.5 s of inserted silence (.21 false fact and.21 hazard against.12 Neutral) — one model, one pause setting, not reproduced in PPlex or in the natural-pause arm.
  • The one positive content effect in Result 1 (word search in PPlex) fails the usefulness test in Result 4.
  • The behaviour is not a threshold artifact (Result 2), not an opportunity artifact and not a permission artifact (Result 3), and not a floor-access artifact (Result 4).

The authors leave the mechanism explicitly open: whether the models fail to notice the problem, or notice it and remain silent, is undetermined. The token-probability probe argues against a noticed-but-suppressed urge in the emission channel, but it reads the output distribution, not comprehension.

Contrast with the "speak when needed" claim#

Full-Duplex Interaction lists proactive interjection — "interrupt when I say something wrong" — as one of the interaction modes that stop being a special harness and become a special case of model behaviour, and Interaction Models carries the same claim. That is Thinking Machines Lab's claim about TML-Interaction-Small, vendor-claim tier, on a research preview, and the wiki has never had a measurement of it.

This paper is the measurement, and it is negative — but on different systems. The two do not meet:

TML's claimPeng et al.
SystemsTML-Interaction-Small, unreleasedMoshi (+Moshika-RL), PersonaPlex (+PPlex-RL), NVIDIA VoiceChat-11B, MiniCPM-o 4.5, Raon-SpeechChat
Tiervendor-claim, first party, self-defined benchmarks (CueSpeak/TimeSpeak)empirical, third party, open models, matched-context protocol
What is shownthe capability is demonstrated in a previewthe capability is absent in every open system measured

Neither supersedes the other. Nothing on this page licenses striking TML's claim, and nothing in TML's claim weakens this result for the seven configurations it covers. What changes is the default: "proactive interjection" can no longer be treated as a property that full-duplex architecture confers. Five open families with the scheduling property do not have the behaviour, one of them explicitly post-trained for interactivity, and two of them refuse to interrupt at all. The burden now sits with the party asserting the capability, and the test exists — this protocol is public (stimuli, model outputs and evaluation code at github.com/vocaliodmiku/take-the-floor) and TML-Interaction-Small has not been run on it. CueSpeak, note, does not substitute: it measures speaking at the right moment with the semantically correct response to a standing user instruction, which is the addressed case this paper finds already works.

The same failure in a different modality#

The structure here is close to the text-agent result on Agent Epistemic Vigilance: models that abstractly know sources have incentives, and still do not act on that knowledge unprompted. Anthropic's Frontier Red Team calls the deficit dispositional rather than cognitive, and this paper is the speech-modality instance of the same shape — the models are not asked to notice a false claim and they do not.

One difference sharpens rather than softens it. The multi-agent listener is never told a source might be unreliable; nothing authorises suspicion, which is part of why the finding is about disposition. Here the authors did authorise it — "if I say anything that is wrong, please just interrupt me straight away" — and onset moved by at most.03. Permission was granted and not used. That is a stronger version of the same negative result than the unprompted condition can produce, and it is the reason to read the two findings as one phenomenon across modalities rather than two coincidences.

The commercial gradient runs the other way (September 2026)#

Worth recording because it is a negative finding about a vendor's aims, not its results. OpenAI's GPT-Live-1 API launch (2026-09-10, vendor-claim) is the corpus's first shipping-product statement of what a voice model's proactivity is for, and it claims nothing whatsoever on the self-selection axis. Every proactivity-adjacent claim in the post is about handling speech the user initiates, or about not speaking:

  • "Smooth interruption handling" and responding "to interruptions and acknowledgements as they happen" — reacting to a user-initiated interruption, the Addressed/Yield side of this page's taxonomy.
  • The Full Duplex Bench v1.5 Interactivity card OpenAI leads with (80.10%) "tests reactions to background speech, speech to another person, listener backchannels, and interruptions" — four reactive conditions, none of them self-selection.
  • "Silent context management": better handling of background noise and silence "without interrupting the conversation or narrating every step out loud" — explicitly a claim about restraint.
  • The customer result OpenAI chose as its lead proof point is a reduction in the model speaking: Speak reports GPT-Live-1 "gave learners more time to think before the language tutor responded, cutting interruptions by almost 80% versus previous turn-based systems."

So the product incentive on a commercial voice layer runs toward speaking less unbidden, and the flagship early-adopter metric is fewer interruptions. That does not bear on whether the capability exists — GPT-Live-1 has never been run through this page's protocol, and the protocol is public. It bears on how the absence should be read: a lab optimising for a tutor that waits is not a lab that would notice a model which never corrects a false claim. The negative result and the product direction are pointed the same way, which makes "nobody built it" at least as live a hypothesis as "nobody can."

A second vendor, the same week, and the same silence on this axis. Google's Gemini 3.8 Live launch (2026-09-15, vendor-claim) makes no self-selection claim either, and its two proactivity-adjacent features are both about the model narrating its own state rather than intervening on the user's content: "early verbal cues like 'Let me check that…' to acknowledge prompts naturally," and "live progress narration to walk users through multi-step background tasks as they progress." Acknowledging a prompt is the Addressed case; narrating a task the user asked for is the Addressed case stretched over time. Neither is the model deciding that something said warrants a turn. So two frontier voice launches five days apart, between them the best-resourced live-dialogue products in existence, and neither claims the capability this page finds absent — which strengthens the "nobody built it" reading considerably, since the marketing is where a capability would surface first if it existed.

One partial exception is worth logging rather than overstating. Google claims the model "automatically detects and transitions between 97 supported languages mid-conversation" — an unprompted, model-initiated change of behaviour triggered by what the user said, with no standing instruction. That is the nearest thing to self-selection any vendor has claimed, and it is a decision about the channel rather than about the content: switching languages serves the user's evident intent, where correcting a false claim contradicts it. The asymmetry is the interesting part — unprompted accommodation ships; unprompted contradiction does not.

The silent model's own technical report, and the same fact valued in reverse (September 2026)#

The vault now holds the first-party report on one of the seven configurations. NemotronLabs VoiceChat (NVIDIA, arXiv 2609.21967, 2026-09-18, empirical) is the technical report for NVIDIA-NemotronLabs-VoiceChat-11B — the same Hugging Face checkpoint Peng et al. cite as their reference [3], published the day after their measurement. The identification is exact, not inferred, and it produces the most useful comparison on this page: the same behaviour, scored by two parties, in opposite directions.

  • Peng et al. score it as a floor. VoiceChat is .00 on nine of ten conditions and.20 on inserted silence: it will not take the floor while the user is speaking, under any trigger, including an explicit hazard and an explicit instruction to interrupt. A model that will not interrupt cannot intervene.
  • NVIDIA scores the identical property as its headline result. On Full-Duplex-Bench 1.0 the same checkpoint posts the lowest pause-handling takeover rate of any open-weight system in its table — 15.3% synthetic, 25.5% CANDOR — against Moshi's 98.5 / 98.0, Freeze-Omni's 64.2 / 48.1 and PersonaPlex's 35.8 / 43.1. On that benchmark lower is unambiguously better, because the pause track exists to catch a model that grabs the floor during a within-turn hesitation.

These are the same disposition. FDB 1.0's pause track rewards silence during ongoing user speech; Peng et al.'s protocol asks what it costs. Nothing is contradicted and the page's thesis is sharpened: the benchmark the field optimizes for pays for restraint and prices intervention at zero. A system engineered to the top of the pause leaderboard is, by construction, being trained toward the corner of the space where content-driven intervention is impossible — and the two labs never have to disagree, because no instrument scores both at once. (Note the shape of the trade in the same table: Moshi, which takes the floor 98.5% of the time during pauses, is also the configuration Peng et al. find most willing to speak mid-turn among the Moshi/PersonaPlex comparison set. The pause TOR column and the onset column are close to inverses.)

The report also supplies the mechanism, which Peng et al. could only probe from outside. Onset is a trained frame-level target: on the 80 ms agent-text timeline, BOS marks response onset and EOS marks the stop, with everything else padding, and the SFT loss weights those tokens at 12.5 for begin-of-turn and 7.5 for end-of-turn against 1.0 for padding — so when to start and stop speaking is upweighted an order of magnitude over what to say. And §4 adds an override the black-box probe could not see: when the model "fails to natively handle turn-taking/barge-in scenarios," a heuristic endpointer built on user speech/silence activity forcibly injects BOS/EOS tokens into the model. So the floor policy Peng et al. measured is partly learned from a speech/silence-shaped objective and partly a hand-written activity heuristic sitting on top of it — neither of which has any channel through which the content of what the user is saying could bear on the decision to speak. That is a structural explanation for a.00, not just a measurement of one.

Two limits on how far to carry this. The FDB numbers are NVIDIA scoring NVIDIA on a single run with no error bars, and every baseline row in that table is copied from third-party reports rather than re-run. And the report says nothing about proactivity, intervention or interruption-worthiness anywhere in its 19 pages — the silence Peng et al. measured is not a design goal the paper defends, it is a property nobody at either lab framed as a choice.

A benchmark that scores restraint directly, and every model fails it hardest (September 2026)#

This page's standing account is that the field's instruments pay for restraint and price intervention at zero. OmniVChat (He, Chu, Chen et al., Alibaba Qwen Team + CUHK + SJTU, arXiv 2609.21465, 2026-09-18, empirical) is the first source in the corpus where restraint is not a by-product of the metric but an explicit graded criterion — and the result runs opposite to everything above.

The construct. Two paired subcategories under Voice Turn-Taking, identical in surface conditions and differing only in whether the user finished:

  • UPC — "User Paused but not Completed." The user stops mid-request amid visual distraction. The rubric's correct reply is to wait or give a brief acknowledgment. Answering is the failure.
  • UCR — "User Completed Request." Same distraction, request finished. Answering is now correct.

The pair is the discriminator this page has been asking for in a different key: it separates cannot read the completion state from will not hold the floor, because a model that scores well on UCR and badly on UPC is not silent-by-disposition — it is speaking when it should not.

The result. Across the six judges in the paper's sensitivity study, all six rank DSLP-VTT-UPC as the weakest of the seventeen subcategories, on 2,800 replies from a frontier commercial model. It is the floor of the benchmark, and it is the one subcategory whose correct answer is don't take the turn yet.

Why this does not contradict the page and does sharpen it. Peng et al.'s finding is about self-selection — no one addressed the model, nothing authorised a turn, and it stays silent. UPC is the Yield case in that taxonomy inverted: the user is mid-turn and has not yielded, and the model takes the floor anyway. Both failures are the same missing competence read from opposite sides — models are not tracking turn state, they are tracking whether a question-shaped thing has arrived. A model that speaks the instant a request sounds complete and stays silent when a hazard is complete is running on surface form in both directions, which is precisely what Peng et al.'s Rhetorical and Turn-end surface-form controls were built to detect.

What it does not establish. This is offline audio-visual comprehension on pre-rendered clips, not a duplex system with a live floor: there is no latency measurement, no barge-in, and the recorded probe "does not test live interruption." The clips are also engineered against the failure they measure — the generator forbids a mid-dialogue probe from explicitly reporting a dropped line or using the word again, so the cue is genuinely prosodic and situational rather than lexical, but the distribution is designed rather than observed. And UPC's floor is a rubric score under a text-only judge that never hears the pause, not a measured onset time. The instrument and its COI are catalogued on Interactivity Benchmarks.

Open Questions#

  • Do the models fail to notice the false claim or the hazard, or notice and stay silent? The text-token probe reads the emission distribution, not comprehension, so it cannot separate them — and the two failures have opposite fixes (better grounding vs. a changed speaking policy). Falsifiable on the released stimuli: probe the internal representation, or hand the same transcript to the same model as a text question and check whether it flags the claim it declined to challenge in speech.
  • Is the absence a pretraining-data property or a post-training one? Other-correction is dispreferred in human conversation, so dialogue corpora are thin in exactly this behaviour; but both interactivity-aligned RL arms here lower content-cue onset below their own baselines, which is a post-training effect pointing the same way. Settled by an RL arm that rewards content-triggered onset on matched stimuli and reports whether turn-taking quality survives it.
  • Do delegating architectures — a duplex frontend with a text backend holding the content understanding — intervene on content better than duplex-native models? None of the seven configurations delegates, and the frontend-backend systems in this corpus are evaluated on tool calls, never on uninvited correction. The experiment is cheap: run this protocol against a frontend-backend system and against a cascaded production voice stack.

Connections#

  • Native Multimodal Modeling: Fusion Depth and I/O Duality — surfaces two streaming instruments built for this construct: ThinkStream's Watch-Think-Speak (scoring response timing, not only accuracy) and AURA's proactive QA (responding when an event occurs with no explicit query)
  • Full-Duplex Interaction — the scheduling property; this page is the decision to use it, and the measurement its proactive-interjection bullet lacked
  • Interactivity Benchmarks — where the protocol sits as an instrument alongside FD-bench, CueSpeak and the tool-calling cluster
  • Interaction Models — carries the "interactivity as model behaviour" claim this result bounds
  • Agent Epistemic Vigilance — the same disposition failure measured in text multi-agent systems; here with permission explicitly granted
  • LLM-as-a-Judge — the intervention rates are a single unvalidated judge's verdicts
  • Time-Aligned Micro-Turns — the mechanism that makes acting during the user's turn possible at all; this paper shows the mechanism does not supply the reason
  • NVIDIA — VoiceChat-11B is theirs, and it is one of the two configurations that will not interrupt ongoing speech, and its own technical report scores that silence as a first-place pause-handling result
  • GPT-Live — production audio full-duplex, untested by this protocol; and, from September 2026, a shipping API whose every proactivity claim is about handling interruptions or about restraint
  • Gemini 3.8 Live — the second shipping voice product to claim nothing on this axis, and the one that claims unprompted language switching instead

Sources#

  • Full-Duplex Speech Models Take the Floor When Asked, Not When Needed — Peng, Nuchged, Fu & Yao (University of Connecticut / UT Austin / X Square Robot / University of Oxford), arXiv 2609.19596, 2026-09-17 (empirical, 5pp, ICASSP-format). Parse note: PDF-derived (docling: 2.126.0, MLX layout/table stages); ingest verify was clean on all five checks. All four tables were nonetheless re-read cell-for-cell from pdftotext -layout and every value matches digit-for-digit, including Table 2's spanning Addressed/Self-select/Yield group header, which docling flattens by repeating the group label once per spanned column. Figures 2 and 3 were viewed: Fig. 2's per-panel onset labels (.07/.34/.83,.08/.15/.12,.41/.35/.68) reproduce Table 2 exactly, and Fig. 2 is the evidence that Raon's high rate is near-continuous speech rather than triggered response. A stray affiliation-footnote fragment ("3 X Square Robot") is welded into §1's closing prose in the docling body; cosmetic, no number affected.
  • Interaction Models: A Scalable Approach to Human-AI Collaboration — Thinking Machines Lab, 2026-05 (vendor-claim): the proactive-interjection claim this result is contrasted against
  • Build more natural voice experiences with GPT‑Live‑1 in the API — OpenAI, 2026-09-10 (vendor-claim): read here only for what it claims — interruption handling, the FDB v1.5 Interactivity card's four reactive conditions, "silent context management," and Speak's ~80% interruption reduction. It makes no self-selection claim and supplies no measurement on this page's axis
  • NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities — Balam, Bartley, Casanova et al. (NVIDIA), arXiv 2609.21967, 2026-09-18 (empirical, 19 pp): the first-party technical report on NVIDIA-NemotronLabs-VoiceChat-11B, the checkpoint Peng et al. cite as reference [3] and score.00 on nine of ten conditions. Read here for Table 1 (FDB 1.0 pause TOR 15.3 / 25.5, lowest open-weight), §2.1's frame-level BOS/EOS turn-taking targets, §3.2's 12.5 / 7.5 / 1.0 loss weighting of begin-of-turn / end-of-turn / padding, and §4's heuristic endpointing fallback that forcibly injects BOS/EOS. NVIDIA scoring NVIDIA, single run, every baseline row copied from third-party reports. Full treatment on Full-Duplex Interaction and Interaction / Background Model Split
  • Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking — Google (Gemini Audio Team), 2026-09-15 (vendor-claim): read here only for what it claims — early verbal cues, progress narration and automatic mid-conversation language switching. It makes no self-selection claim and supplies no measurement on this page's axis
§ end
Cited by 11
  • Full-Duplex Interaction×7

    The consequence for this page: the modes listed above are what the architecture makes possible, and…

  • Interactivity Benchmarks×5

    The hardest subcategory is knowing when not to speak. All six judges in the sensitivity study rank…

  • Agent Epistemic Vigilance×3

    So the asymmetry this page names — machinery for reading a source, none for discounting one, and no…

  • Time-Aligned Micro-Turns×3

    (Qualified 2026-09-21 by Peng et al., empirical. The mechanism removes the constraint on acting…

  • Interaction Models×2

    Content Driven Intervention — the "speak when needed" mode, measured on five open families and…

  • NVIDIA×2

    Three things make this NVIDIA's most disclosure-heavy publication in the corpus. It is open-weight…

  • Interaction / Background Model Split

    Open source, unlike its sibling. Hu et al. released no frontend weights. This checkpoint is on…

  • LLM-as-a-Judge

    Content Driven Intervention — the opposite end of the validation spectrum from the specimens above:…

  • Interaction & Multimodal

    Content Driven Intervention — Speaking because the content warrants it — correcting a false claim,…

  • Native Multimodal Modeling: Fusion Depth and I/O Duality

    The full-duplex rows corroborate the vault's own numbers from a neutral source: Moshi Eval at a 200…

  • Open Questions Backlog

    Content Driven Intervention ×3 (oldest 7d) — Do the models fail to notice the false claim or the…

Related articles