高级检索

融合资源丰富语言信息的藏文词向量表示方法

A Tibetan Word Embedding Representation Method Incorporating High Resource Language Information

  • 摘要: 藏文词向量表示是基于深度学习的藏语自然语言处理的基础任务。针对藏文作为低资源语言,仅依赖藏文单语资源难以充分捕捉词深层语义信息的问题,该文提出了一种融合资源丰富语言信息的藏文词向量表示方法(TiWE-RL),该方法创新性地引入并融合资源丰富语言信息,以补充和增强藏文词向量的语义表征能力。经实验验证融合资源丰富汉、英语言信息的藏文词向量表示方法(TiWE-RL)在藏文词相似度和相关性评测集TWordSim215和TWordRel215上的得分分别为58.77和54.24,比现有最佳基线模型TCCWE分别提升了7.75和3.90个百分点。特别地,融合藏文本体信息和资源丰富语言信息的双模型协同架构模型TiWE-RL-TCCWE既能通过TiWE-RL模型有效融合汉英文的语义表征能力,又能利用TCCWE模型中藏文字符以及其构件特征捕捉藏文语义,显著提升模型的整体性能,在藏文词相似度和相关性评测集TWordSim215和TWordRel215上的得分分别为59.68和57.51,比现有最佳基线模型TCCWE分别提升了8.66和7.17个百分点。该研究有效缓解了藏文因语料稀疏导致的语义表征不足问题,为低资源语言词向量表示提供了新思路。

     

    Abstract: Tibetan word embedding representation is a fundamental task for Tibetan Natural Language Processing. To address the challenge of insufficient semantic representation caused by the limited monolingual resources in Tibetan as a low-resource language, this study proposes a novel Tibetan word embedding representation method (TiWE-RL), which enhances the Tibetan semantic representation by integrating high-resource language information such as Chinese and English embeddings.Experimental results demonstrate that TiWE-RL achieves scores of 58.77 and 54.24 on the Tibetan word similarity and relevance evaluation sets (TWordSim215 and TWordRel215), outperforming the state-of-the-art baseline model TCCWE by 7.75% and 3.90%. Specifically, a dual-model collaborative framework (TiWE-RL-TCCWE) is proposed, which combines the high-resource language information with the Tibetan character and component features, sourced from TiWE-RL and TCCWE. This framework achieves scores of 59.68 and 57.51 on TWordSim215 and TWordRel215, surpassing TCCWE by 8.66% and 7.17%, respectively.

     

/

返回文章
返回