高级检索

基于语义协同的多模态知识图谱嵌入

Multimodal Knowledge Graph Embedding Based on Semantic Collaboration

  • 摘要: 充分利用多模态信息对增强知识图谱嵌入具有重要意义,尤其是在不同语义环境下,实体对于多模态信息的语义偏好可能存在显著差异。为此,该文提出了一种基于图像语义协同的多模态知识图谱嵌入模型,充分挖掘全局与局部两种互补的图像语义特征,以提升知识嵌入性能。其中,全局语义模块通过评估实体图像的视觉一致性提取全局语义特征,而局部语义模块则结合三元组的上下文语境,分析图像在特定语境中的语义差异。知识表示模块动态整合文本与图像数据,并采用基于复数空间嵌入的解码器,进一步优化跨模态表示。实验结果表明,该模型在多个公开数据集上的关键性能指标相比基线方法提升了1~4个百分点,展示了其在知识图谱补全任务中的有效性,同时也体现了在复杂语义场景下链接预测的稳定性。

     

    Abstract: Effectively utilizing multimodal information to enhance knowledge graph embeddings is crucial, as entities may exhibit varying semantic preferences for multimodal information in different contexts. This paper proposes a multimodal knowledge graph embedding model that leverages both global and local image semantic features to improve embedding performance. The global semantic module captures visual consistency to extract global features, while the local semantic module analyzes contextual semantic variations within triples. The model dynamically integrates textual and visual data and employs a complex space-based decoder to optimize cross-modal representations. Experimental results on multiple public datasets show 1~4% improvement over the baselines.

     

/

返回文章
返回