首页 > AI前沿 > Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios

Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios

arXiv自然语言 2026-03-09 19:22 5 阅读 查看原文

Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but remains out of reach for most under-resourced languages due to the lack of labeled video corpora for training.

Synthetic visual data have been shown to be an effective augmentation strategy for addressing AV data scarcity. However, a more challenging scenario arises for languages such as Catalan, where no real audiovisual data are available for training.

In this Study

In this study, we investigate whether AVSR can be bootstrapped in such a zero-AV-resource setting, using synthetic visual data as the sole source of visual supervision.

We synthesize over 700 hours of talking-head video and fine-tune a pre-trained AV-HuBERT model.

On a manually annotated Catalan benchmark, our model achieves near state-of-the-art (SOTA) performance with much fewer parameters and training data than SOTA ASR systems such as Whisper-large-v3.

Our model outperforms an identically trained audio-only baseline, and preserves multimodal advantages under acoustic degradation.

Scalable synthetic video thus offers a viable substitute for real recordings in zero-AV-resource AVSR.