PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 30, 2026Electronics0 citationsOpen Access

Bridging Cross-Modal Semantic Gaps with Multi-Source Semantic Anchors in Knowledge-Based Visual Question Answering

View Full Paper
HJHu JunMingJZJinxiong ZhangFZFeng Zhan

Key Points

  • The aim is to improve knowledge-based visual question answering by addressing cross-modal semantic gaps.
  • Developed a unified framework using multi-source semantic anchors for aligning vision and textual features.
  • Implemented a cross-residual gating mechanism to reduce modality noise during cross-modal fusion.
  • Adopted contrastive learning to enhance cross-modal alignment and used a retrieve-then-read pipeline for answer generation.
  • Framework demonstrates improved performance on OK-VQA, FVQA, and A-OKVQA datasets.
  • Outperforms state-of-the-art methods across multiple evaluation metrics.

Abstract

Knowledge-based visual question answering (KB-VQA) requires leveraging external knowledge relevant to the image to assist reasoning. Existing methods typically convert images into a single textual description for knowledge retrieval or directly rely on the implicit knowledge within large language models to generate answers. However, a single textual description struggles to preserve fine-grained visual information such as object attributes and scene text, limiting retrieval quality. Meanwhile, naively fusing multi-source information tends to introduce modality noise, undermining reasoning accuracy. To address these issues, we propose a unified framework that constructs multi-source semantic anchors to bridge the cross-modal semantic gaps among vision, questions, and external knowledge. Specifically, we unify image captions, object tags, and optical character recognition (OCR) text as semantic anchors. These anchors serve as shared intermediaries to pre-align visual and textual features, avoiding direct interaction between heterogeneous modalities. During cross-modal fusion, a cross-residual gating mechanism adaptively suppresses modality noise by leveraging the semantic anchors as stable references. The framework further integrates contrastive learning to strengthen cross-modal alignment and employs a retrieve-then-read pipeline for open-domain answer reasoning. Experiments on the OK-VQA, FVQA, and A-OKVQA datasets demonstrate that the proposed framework outperforms state-of-the-art methods across multiple metrics, validating the effectiveness and robustness of the proposed framework.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

JunMing et al. (2026) studied this question.

synapsesocial.com/papers/69f2a42a8c0f03fd67763206https://doi.org/10.3390/electronics15091837
Ask AI
Helpful
Bookmark
Share
View Full Paper