首页 > AI前沿 > Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

Beyond Speech Captions: Speech-Rewarded Style Planning for Conversational Text-to-Speech

arXiv自然语言 2026-10-08 16:12 4 阅读 查看原文

Natural-language style descriptions provide an interpretable interface between large language models (LLMs) and controllable text-to-speech (TTS). However, using descriptions as pseudo-labels compresses target acoustics into text, and descriptive fidelity need not imply effective control of a particular synthesizer.

We empirically show that speech-text alignment only weakly predicts downstream acoustic similarity among candidate instructions for the same utterance. We therefore propose Speech-Rewarded Style Planning (SRSP), which trains a text-based style planner through a frozen downstream TTS model.

Given dialogue history and response text, the planner generates candidate instructions and is optimized with group-relative policy optimization (GRPO), using the teacher-forced likelihood of target speech tokens as the reward.

On an English subset of the ISCSLP 2026 CoT-TTS corpus, SRSP achieves higher speech-style and emotion similarity to target speech and lower mel-cepstral distortion than the Base LLM and target-audio-informed captioning baselines.

LLM-based expressive speech evaluation further shows gains over all baselines in contextual appropriateness and reference consistency.