Self-supervised speech encoders contain linguistic and paralinguistic information in a shared, entangled representation space.
We combine a TopK sparse autoencoder with route-specific supervision and cross-factor adversaries.
Experiments
Across frozen SPEAR and WavLM encoders, independent probes show factor-specific retention and suppression:
linguistic information remains stronger in the linguistic route, while paralinguistic factors, including speaker identity, emotion, and prosody, are retained in the paralinguistic route and substantially reduced in the linguistic route.
Route Organisation
The route organisation learned on LibriSpeech persists on MSP-Podcast without representation-side retraining.
Feature-space Route Interventions
Feature-space route interventions further transfer the swapped factor while largely preserving the information carried by the unchanged route.
Conclusions
These results show consistent route-selective separation across encoders, corpora, independent probes, and representation-level interventions.