首页 > AI前沿 > Full-Duplex Speech Models Take the Floor When Asked, Not When Needed

Full-Duplex Speech Models Take the Floor When Asked, Not When Needed

arXiv自然语言 2026-09-17 10:21 1 阅读 查看原文

Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak.

Human listeners speak when addressed or when the speaker stops, but also self-select to correct a false claim, supply a missing word, or warn of danger. We ask whether full-duplex models do the same.

To separate the reason to speak from the opportunity, we construct context-matched English monologues in which only the trigger utterance varies within a topic, define 10 conditions from turn-allocation rules, and compress inter-word pauses to limit opportunities created by silence.

Across five model families, being addressed and silence are far more reliable triggers than false facts or hazards.

Frame-level text-token probabilities in Moshi and PersonaPlex are lower for false facts than for Neutral when averaged over the first 2 s after trigger end. Pauses or permission to interrupt do not close this gap either.

Given the floor, Moshi and PersonaPlex answer most direct questions, yet the proportion of non-empty false-fact replies that challenge the claim is only .14--.15, and the proportion of hazard replies that warn of danger is .04--.07.

This paper thus identifies a gap in both speech initiation and response content. Closing it requires genuine content understanding and intervention decisions grounded in it.