Abstract:
Effectively utilizing multimodal information to enhance knowledge graph embeddings is crucial, as entities may exhibit varying semantic preferences for multimodal information in different contexts. This paper proposes a multimodal knowledge graph embedding model that leverages both global and local image semantic features to improve embedding performance. The global semantic module captures visual consistency to extract global features, while the local semantic module analyzes contextual semantic variations within triples. The model dynamically integrates textual and visual data and employs a complex space-based decoder to optimize cross-modal representations. Experimental results on multiple public datasets show 1~4% improvement over the baselines.