Advanced Search
KAN Taiji, CAI Zhijie. A Tibetan Word Embedding Representation Method Incorporating High Resource Language InformationJ. Journal of Chinese Information Processing, 2026, 40(6): 63-70. DOI: 10.3969/j.issn.1003-0077.2026.06.007
Citation: KAN Taiji, CAI Zhijie. A Tibetan Word Embedding Representation Method Incorporating High Resource Language InformationJ. Journal of Chinese Information Processing, 2026, 40(6): 63-70. DOI: 10.3969/j.issn.1003-0077.2026.06.007

A Tibetan Word Embedding Representation Method Incorporating High Resource Language Information

  • Tibetan word embedding representation is a fundamental task for Tibetan Natural Language Processing. To address the challenge of insufficient semantic representation caused by the limited monolingual resources in Tibetan as a low-resource language, this study proposes a novel Tibetan word embedding representation method (TiWE-RL), which enhances the Tibetan semantic representation by integrating high-resource language information such as Chinese and English embeddings.Experimental results demonstrate that TiWE-RL achieves scores of 58.77 and 54.24 on the Tibetan word similarity and relevance evaluation sets (TWordSim215 and TWordRel215), outperforming the state-of-the-art baseline model TCCWE by 7.75% and 3.90%. Specifically, a dual-model collaborative framework (TiWE-RL-TCCWE) is proposed, which combines the high-resource language information with the Tibetan character and component features, sourced from TiWE-RL and TCCWE. This framework achieves scores of 59.68 and 57.51 on TWordSim215 and TWordRel215, surpassing TCCWE by 8.66% and 7.17%, respectively.
  • loading

Catalog

    Turn off MathJax
    Article Contents

    /

    DownLoad:  Full-Size Img  PowerPoint
    Return
    Return