高级检索

面向大模型的最优藏文字表构建研究

Construction of an Optimal Tibetan Character Set for Large Language Models

  • 摘要: 字表构建在语言模型中充当连接原始文本和模型输入的桥梁,是大语言模型和许多其他智能信息处理的先决条件,但目前还没有公开的藏文字表。基于此,该文旨在分析大数据建模中表现最优的藏文字表,以及是否能够在挖掘文字内部的细粒度特点下提出一种简单高效的藏文字表构建方法。首先通过经典藏文文法理论,全面分析藏文字符、字丁、字、块和词的概念及特点,接着基于这些特征定义了藏文字表,整理了藏文字表数据集,并结合大模型的词元切分特征提出了一种新颖的藏文字表构建方法,对比了该方法和传统方法。经实验表明,Bod-WL藏文字表在传统的藏文字或音节处理中使词汇量减少了近28.67%,同时大模型的三种主流词元切分算法Bod-BPE、Bod-WordPiece和Bod-Unigram在语言建模性能上分别提升了2.01、2.87和2.96个百分点。为助力藏语智能信息领域的学者开展深度研究,笔者将会开源所构建的藏文字表及其资源

     

    Abstract: Character sets bridge the gap between raw text and model inputs in language models and serve as fundamental components of large language models and other intelligent information processing systems. However, no public Tibetan character set is currently available. This study aims to identify an optimal Tibetan character set for large-scale data modeling and develop a simple and efficient construction method that captures the fine-grained linguistic features of the Tibetan writing system. Based on classical Tibetan grammar theory, we analyze Tibetan characters, stacks, syllables, blocks, and words and establish a formal definition of the Tibetan character set. A Tibetan character set dataset is then compiled according to this definition. Considering the characteristics of tokenization algorithms used in large language models, we propose a novel construction method and compare it with traditional approaches. Experimental results show that the Bod-WL character set reduces vocabulary size by 28.67% compared with conventional character-level and syllable-level processing methods. Relative to the baseline, the Bod-BPE, Bod-WordPiece, and Bod-Unigram tokenization algorithms improve language modeling performance by 2.01, 2.87, and 2.96 percentage points, respectively. These results demonstrate the effectiveness of the proposed method for Tibetan language modeling. The constructed Tibetan character set and related resources are publicly available at https://github.com/AI-Bod/.

     

/

返回文章
返回