首页 > AI前沿 > Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters

arXiv自然语言 2026-10-05 06:25 5 阅读 查看原文

Much of expressive text-to-speech research rests on an untested assumption that written text carries enough information to select an appropriate prosodic style for its delivery.

Text-predicted style models improve listener preference, and expressive-appropriateness evaluation presupposes that context constrains style, yet neither measures the assumption itself.

This paper tests it as a falsifiable hypothesis against style labels derived from acoustics alone.

For each of six speakers in a 1,200-hour conversational corpus, utterances are clustered in the spaces of five speech models, including a prosody-only control, and the cluster of held-out utterances is predicted from twelve text embedding models.

Three controls are applied: utterance length is erased from the speech embeddings; accuracy is scored against the majority-class floor of unbalanced clusters rather than uniform chance; and a bag-of-words baseline measures word identity alone.

Text predicts the cluster above that floor for all six speakers (+0.111 top-3 accuracy), but bag-of-words achieves three quarters of this.

Sentence embeddings add only +0.026, largest for encoders not trained for sentence semantics and reversed by tree-based probes for all others.

Acoustic clusters are not compact in text embedding space in any of 360 configurations.

The prosody-only space weakens the association for five speakers, but not for the speaker showing it most strongly.

Text thus informs these delivery clusters mainly through word choice, whether as a cue to prosody or as a marker of topic and recording situation, and reference-free style selection cannot assume more.