高级检索

基于相对位置编码注意力和双向解码器的增强型E-Branchformer语音识别模型

Enhanced E-Branchformer Speech Recognition Model Based on Relative Positional Attention and Bidirectional Decoding

  • 摘要: 在端到端语音识别领域,E-Branchformer模型已在多个任务上取得了较好性能。为评估并进一步提升其在低资源语言场景下的表现,该文提出一种增强型E-Branchformer语音识别模型。该模型利用基于相对位置编码的缩放点积注意力机制对传统E-Branchformer模型编码器的自注意力模块进行优化,以增强对序列位置信息以及局部和全局依赖关系的捕捉能力;同时采用混合精度训练策略,以提高训练效率并降低显存占用;采用双向解码器和CTC/Attention联合解码策略,以学习到更丰富的上下文信息。实验表明,该文提出的增强型E-Branchformer模型相较于传统E-Branchformer模型在低资源的藏语语音识别数据集上字错率(Character Error Rate,CER)平均降低了0.76%;在加噪后的藏语语音识别数据集上较传统E-Branchformer模型平均降低了4.22%;在开源的Aishell-1数据集上字错率平均降低了0.08%。实验结果显示出增强型E-Branchformer模型在提升藏语和汉语语音识别性能的有效性和在加噪环境下的鲁棒性。增强型E-Branchformer模型通过架构优化和混合精度计算策略,在精度、效率和内存占用之间实现了良好平衡。

     

    Abstract: In the field of end-to-end speech recognition, the E-Branchformer model has achieved competitive performance across a variety of tasks. To evaluate and further improve its effectiveness in low-resource language scenarios, this paper proposes an enhanced E-Branchformer speech recognition model. The proposed model optimizes the self‑attention module of the conventional E‑Branchformer encoder by employing a scaled dot‑product attention mechanism with relative position encoding, thereby strengthening the ability to capture sequential positional information as well as local and global dependencies. In addition, mixed‑precision training is adopted to improve training efficiency and reduce GPU memory footprint. Furthermore, a bidirectional decoder together with a joint CTC/attention decoding strategy is utilized to learn richer contextual information. Experimental results demonstrate that the proposed enhanced E-Branchformer model achieves average Character Error Rate (CER) reductions of 0.76% on the low-resource Tibetan speech recognition dataset, 4.22% on a noisy Tibetan speech dataset, and 0.08% on the open-source Aishell-1 Mandarin benchmark, compared to the conventional E-Branchformer baseline. These findings validate the effectiveness of the enhanced model in improving both Tibetan and Mandarin speech recognition performance, as well as its robustness under noisy acoustic conditions. Furthermore, through architectural refinements and the integration of mixed-precision training strategies, the enhanced E-Branchformer attains a favorable trade-off among recognition accuracy, computational efficiency, and memory footprint.

     

/

返回文章
返回