高级检索

基于自提示多模态大语言模型和语义感知离散扩散模型的图像描述生成算法

Exploring Semantic-aware Discrete Diffusion Model via Self-prompting Multimodal LLMs for Image Captioning

  • 摘要: 近年来,非自回归图像描述生成技术凭借其双向传播和并行词语生成的能力受到广泛关注。与此同时,基于离散扩散方法的研究也取得了显著进展。然而,在离散噪声添加与去除过程中,现有方法仍面临图像文本关联性低、目标物体遗漏、描述准确性不足以及词语重复等关键问题。为应对这些挑战,该文提出一种基于语义感知的离散扩散模型。该模型通过可学习查询机制构建语义感知模块,以捕捉与图像物体级语义特征的潜在关联,从而更好地生成图像描述。在此基础模型之上,该文进一步引入自提示优化框架,利用大语言模型生成与图像细节内容更相符的丰富描述。在COCO数据集上的综合实验表明,该文方法在图像描述生成任务中取得一定的进展,其性能优于现有的相关方法。

     

    Abstract: In recent years, non-autoregressive image captioning has gained significant attention due to its capabilities in bidirectional propagation and parallel word generation. Meanwhile, considerable progress has been made in research on discrete diffusion-based approaches. However, during the processes of discrete noise addition and denoising, existing methods still face critical challenges such as weak image-text relevance, object omission, inaccurate descriptions, and word repetition. To address these issues, this paper proposes a semantic-aware discrete diffusion model. This model incorporates a learnable query mechanism to construct a semantic perception module, which captures latent correlations with object-level semantic features in images. Building upon this foundational model, we further introduce a self-prompting optimization framework that leverages large language models to generate richer descriptions that better align with image details. Comprehensive experiments on the COCO dataset demonstrate that our method achieves notable improvements in image captioning tasks and outperforms existing approaches.

     

/

返回文章
返回