高级检索

基于检索增强和知识蒸馏的汉越跨语言查询扩展方法

Chinese-Vietnamese Cross-lingual Query Expansion Method Based on Retrieval Augmentation and Knowledge Distillation

  • 摘要: 汉越跨语言查询扩展旨在增加与中文查询语句语义相同的术语和概念,并将扩展后的中文查询语句转换为越南语查询语句。然而,由于大规模语言模型的参数量庞大且计算资源消耗高,难以直接应用于跨语言查询扩展任务。同时,多语言预训练模型在低资源语言场景下的推理和生成能力表现不佳。为了解决这一问题,该文提出了一种基于检索增强和知识蒸馏的汉越跨语言查询扩展方法。通过知识蒸馏和检索增强,把大规模语言模型具备的思维链生成能力和检索得到的外部知识融入小参数规模多语言预训练模型,从而提升其思维链生成能力。实验结果显示,该方法在MLQA、XQuAD公共数据集以及该文构建的汉越跨语言查询扩展数据集上取得的性能优于基线模型,其中,MAP、Recall、NDCG和MRR分别提高了3.4%、1.6%、2.9%和3.4%。

     

    Abstract: Chinese-Vietnamese cross-lingual query expansion aims to add terms and concepts matching the semantics of Chinese queries and convert them into Vietnamese. To deal with low-resource language in a reasonable computational costs, this paper proposes a Chinese-Vietnamese cross-lingual query expansion method based on retrieval augmentation and knowledge distillation. By leveraging knowledge distillation and retrieval augmentation, the method injects the chain-of-thought generation capabilities and external knowledge from large language models into multilingual pre-trained models with fewer parameters. Experimental results on MLQA, XQuAD, and a constructed Chinese-Vietnamese dataset demonstrate that this method outperforms baseline models, with improvements of 3.4%, 1.6%, 2.9%, and 3.4% in MAP, Recall, NDCG, and MRR metrics, respectively.

     

/

返回文章
返回