Advanced Search
HU Jian, DONG Ling, GAO Shengxiang, WANG Wenjun, XIANG Yan, YU Zhengtao. Optimizing Whisper for Low-resource Speech Recognition via Self-supervised Representation DistillationJ. Journal of Chinese Information Processing, 2026, 40(8): 165-174. DOI: 10.3969/j.issn.1003-0077.2026.08.016
Citation: HU Jian, DONG Ling, GAO Shengxiang, WANG Wenjun, XIANG Yan, YU Zhengtao. Optimizing Whisper for Low-resource Speech Recognition via Self-supervised Representation DistillationJ. Journal of Chinese Information Processing, 2026, 40(8): 165-174. DOI: 10.3969/j.issn.1003-0077.2026.08.016

Optimizing Whisper for Low-resource Speech Recognition via Self-supervised Representation Distillation

  • Whisper is a powerful multilingual automatic speech recognition model and performs well on high-resource languages such as English. Its accuracy on some low-resource languages, including Burmese, remained limited because of insufficient pretraining data. This study proposes a low-resource ASR optimization method based on self-supervised representation distillation. A cross-model representation distillation mechanism is used to transfer knowledge from a self-supervised speech model to the Whisper encoder. Experiments on Burmese, Khmer, Uzbek, and Punjabi achieves the character error rate reduction of 4.8%, 6.5%, 6.0%, and 6.3%, respectively.
  • loading

Catalog

    Turn off MathJax
    Article Contents

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return