高级检索

融合细粒度错误分析的机器译文自动评价方法

Automatic Evaluation of Machine Translation Integrated with Fine-grained Error Analysis

  • 摘要: 机器译文自动评价通过量化机器译文与人工参考译文之间的相似性或差异性来实现对机器译文的质量评价,对推动机器翻译发展和应用发挥着重要作用。针对当前主流的基于预训练语言模型的自动评价方法仅输出单一的质量分数、缺乏详细错误说明的不足,以及最新利用大语言模型直接对机器译文进行评分的自动评价方法暴露的分数有偏性问题,该文提出了一种融合细粒度错误分析的机器译文自动评价方法,通过思维链引导大语言模型捕捉机器译文中的具体错误信息,对不同严重程度的错误进行惩罚,并将单位机器译文长度内的错误数量量化为错误分布集中程度,结合错误严重性与错误密度,充分考虑细粒度的错误匹配信息。在WMT’23 机器译文自动评价评测基准数据集上的实验结果表明,与经典的机器译文自动评价方法和参与评测的最优机器译文自动评价方法相比,融合细粒度错误分析的自动评价方法有效提高了其与人类评价的相关性。

     

    Abstract: Automatic evaluation of machine translation can access the quality of machine translation by quantifying the similarity or difference between machine translation and human reference translation. To address the defects in lacking detailed error description and score bias in current pre-trained language model based methods, this paper proposes to integrate fine-grained error analysis for the automatic evaluation method of machine translation. The large language model is guided by the chain-of-thought to capture various errors in machine translation, punish errors of different severity, and quantify the number of errors per unit length of machine translation. Experimental results on the WMT '23 benchmark dataset for automatic evaluation of machine translation show that the automatic evaluation methods integrated with fine-grained error analysis can effectively improve the evaluation results compared with the classical automatic methods and the best results in the evaluation campaign.

     

/

返回文章
返回