高级检索

难负样本增强的对比知识图谱补全

Hard Negative Samples Enhanced Contrastive Knowledge Graph Completion

  • 摘要: 难负样本是影响对比学习性能的关键因素,而现有负采样方法顺带产生的大量易负样本往往会损害知识图谱补全模型的性能。针对对比知识图谱补全问题,该文提出了一种难负样本增强的对比知识图谱补全框架。在该框架下设计了基于文本语义、图谱结构和数据集成三种难负样本生成方法,并由此得到三个有效的难负样本池;然后,分别从三个难负样本池中取出若干负样本加到In-Batch负采样生成的负样本集中,进而形成增强负样本集合;最后,在增强负样本集合和正样本集合上运用对比学习训练对比知识图谱补全模型。在两个公开基准数据集上的实验结果表明,该文所提优化模型在 WN18RR 数据集上的 MRR 和 Hits@1 分别提升 0.1% 和 0.9%,在 FB15k-237 数据集上分别提升 1.8% 和 1.6%。实验结果表明,三种方法均能够产生提升对比知识图谱补全模型性能的难负样本,并且难负样本增强的对比知识图谱补全框架能够充分利用嵌入向量包含的文本语义信息和知识图谱包含的结构信息,达到当前最优的知识图谱补全性能。

     

    Abstract: Hard negative samples are a key factor affecting contrastive learning performance. Aiming at the problem of contrastive knowledge graph completion (CKGC), we propose a hard negative sample enhanced CKGC framework. Under this framework, we design three hard negative sample generation methods based on text semantics, knowledge graph structure, and data ensemble. According to these methods, three effective hard negative sample pools are obtained. Then, several samples are taken from three hard negative sample pools to augment the negative sample set generated by In-Batch negative sampling. Finally, such enhanced negative sample set and the positive samples set are trained through contrastive learning to obtain an optimized CKGC model. Experiments on the WN18RR and the FB15k-237 datasets show that all three methods can generate hard negative samples that improve the performance of the CKGC model.

     

/

返回文章
返回