首页 > AI前沿 > DuplexSpeechBench-Document Grounding: Benchmarking Document Grounding and Hallucinations in Voice Agents

DuplexSpeechBench-Document Grounding: Benchmarking Document Grounding and Hallucinations in Voice Agents

arXiv自然语言 2026-10-02 12:00 6 阅读 查看原文

Voice agents enable low-latency, natural interaction, yet their ability to faithfully ground responses in external documents remains underexplored.

We introduce DuplexSpeechBench-Document Grounding (DSB-DG), a benchmark for evaluating document grounding in voice agents across five professional domains.

DSB-DG targets three failure modes: Context Saturation, which measures grounding under increasing document length; Grounding Decay, which measures retention of document facts across multi-turn dialogue; and Proactive Grounding, which evaluates whether context re-injection mitigates conversational drift.

The benchmark contains 1,636 adversarially verified QA pairs from 50 documents covering five professional domains, and supports fully automatic evaluation of grounding accuracy, hallucination, and response latency.

Across systems spanning cascaded, proprietary full-duplex and real-time, and open-weight speech2speech architectures, we find substantial differences in effective grounding capacity.

While cascaded pipeline (ASR-LLM-TTS) achieves the highest grounding accuracy, Gemini-Live and GPT-Realtime closely trail behind.

Open-weight systems exhibit distinct failure modes, most notably an abrupt context-capacity collapse and multi-turn grounding decay.

More broadly, grounding fidelity degrades with context and conversational load, and failures frequently manifest as unsupported generations rather than abstention.

We show that contextual grounding as a key unresolved challenge for reliable full-duplex voice agents.