Optimizing Whisper for Low-resource Speech Recognition via Self-supervised Representation Distillation
-
Abstract
Whisper is a powerful multilingual automatic speech recognition model and performs well on high-resource languages such as English. Its accuracy on some low-resource languages, including Burmese, remained limited because of insufficient pretraining data. This study proposes a low-resource ASR optimization method based on self-supervised representation distillation. A cross-model representation distillation mechanism is used to transfer knowledge from a self-supervised speech model to the Whisper encoder. Experiments on Burmese, Khmer, Uzbek, and Punjabi achieves the character error rate reduction of 4.8%, 6.5%, 6.0%, and 6.3%, respectively.
-
-