Abstract:
To address the issues of visual modality information loss and multimodal noise in multimodal knowledge reasoning tasks, a knowledge reasoning method based on adversarial networks and relation-guided attention is proposed. Through multimodal adversarial learning, the existing textual modality information is utilized to generate the missing visual modality features, thereby alleviating the information imbalance and entity alignment problems caused by modality loss. Meanwhile, a relation-guided attention mechanism is designed to capture important multimodal features relevant to relation context, thereby reducing multimodal noise to mitigate the problems of misalignment and feature offset duo to incomplete information, feature sparsity, and cross-modal semantic inconsistency. To validate this method, experiments are conducted on the FB15K-237 and WN18RR datasets, comparing the results with 14 methods including RotatE, IMF, and MLSFF, demonstrating the effectiveness of this approach.