高级检索

面向小样本工艺文本实体识别的最近邻提示学习方法

Nearest Neighbor Prompt Learning Model for Few-shot Process Text Entity Recognition

  • 摘要: 该文研究小样本条件下的工艺文本实体识别问题,提出了最近邻提示学习模型(NN-PLM)。该模型利用了工艺文本中字符的构词角色规律来建模外显记忆,通过最近邻搜索为待预测字符检索出与之具有相同或相近构词角色的字符与实体标签对,修正提示学习模型预测概率,改善小样本条件下的识别效果。NN-PLM模型具有与预训练语言模型相同的参数量,无需额外训练。实验结果表明,该文方法在不同规模数据集上的实体识别效果均优于9种对比方法,尤其在5-shot(每类实体5个样本)与10-shot(每类实体10个样本)数据规模下,NN-PLM模型F1值较次优模型分别提升11.88%与7.62%。表明该方法能有效改善因训练数据少导致的模型学习不充分问题,提升了小样本条件下工艺文本的实体识别效果。

     

    Abstract: To address the entity extraction from the process text under few-shot settings, this paper proposes a Nearest Neighbor Prompt Learning Model (NN-PLM). This model utilizes word-building role of characters in the process text to model an explicit memory, and retrieves character-label pairs which have similar word-building role with the predicted characters through a nearest-neighbor search. By correcting the prediction probability of the prompt learning model, the NN-PLM model improves entity recognition performance under few-shot settings. Experimental results demonstrate the effectiveness of NN-PLM compared with nine methods, with 11.88% and 7.62% improvements in F1 in 5-shot and 10-shot settings, respectively.

     

/

返回文章
返回