Accent classifiers are typically trained with a fixed label inventory and cannot accommodate new accent categories as new data becomes available.
Moreover, accented speech corpora often exhibit substantial class imbalance and/or domain shift due to differences in recording conditions across corpora.
We present AccentCL
We present AccentCL, a class-incremental learning framework for English accent classification that is robust to class imbalance and cross-corpus domain shift.
AccentCL extracts multi-layer representations from a frozen Whisper-Large-v3 encoder, optimized with an imbalance-aware cross-entropy loss to reduce bias toward the majority accent classes and a domain mean alignment loss that minimizes distributional mean shift across training corpora.
The label space is then expanded via replay-based continual learning, using the frozen base model for knowledge retention and an old-to-new margin loss to reduce overprediction on newly added classes.
On a five-class accent classification task
On a five-class accent classification task, AccentCL achieves 77.1% balanced accuracy and a 76.9% macro-averaged F1 score.
We further evaluate the model's ability to incrementally incorporate two new accent categories
We further evaluate the model's ability to incrementally incorporate two new accent categories: Spanish-accented and Chinese-accented English.
When adding Spanish-accented English to the pretrained model, AccentCL attains an F1 of 83.3% on the new class while retaining 77.3% balanced accuracy on the base classes.
When subsequently adding Chinese-accented English, it achieves 61.8% F1 on the new class while preserving 77.6% balanced accuracy on the previously learned classes.
These results show that AccentCL enables robust regional accent classification while allowing new accent categories to be added without full retraining.