While predicting prosody from text is an established task in the field, the opposite direction, predicting text that fits a given prosodic pattern, remains largely overlooked.
We find this unfortunate, because this opposite direction could lead to some very interesting use cases.
First Steps in the Prosody-to-Text Direction
Therefore, in this paper, we make the first steps in the prosody-to-text direction by investigating how much of the original sentence can be recovered from its prosodic pattern.
To this end, we fine-tune the Whisper model using only the 12 lowest Mel bins (low-pass filter with approximately 450Hz cutoff), and obtain surprisingly accurate results (WER 36%), with 10% of utterances being recovered perfectly, and 40% of utterances having Word Error Rate at or below 25%.
We also find that, given the correct prefix, the next token was predicted correctly in 79% of cases.
Implications and Future Directions
Our results suggest that the relationship between low-frequency speech features and lexical content is much stronger than previously thought.
Directing more attention to this topic might open the door to new applications, such as using prosody to guide text generation of modern LLMs.