Knowledge-based visual question answering (KB-VQA) requires leveraging external knowledge relevant to the image to assist reasoning. Existing methods typically convert images into a single textual description for knowledge retrieval or directly rely on the implicit knowledge within large language models to generate answers. However, a single textual description struggles to preserve fine-grained visual information such as object attributes and scene text, limiting retrieval quality. Meanwhile, naively fusing multi-source information tends to introduce modality noise, undermining reasoning accuracy. To address these issues, we propose a unified framework that constructs multi-source semantic anchors to bridge the cross-modal semantic gaps among vision, questions, and external knowledge. Specifically, we unify image captions, object tags, and optical character recognition (OCR) text as semantic anchors. These anchors serve as shared intermediaries to pre-align visual and textual features, avoiding direct interaction between heterogeneous modalities. During cross-modal fusion, a cross-residual gating mechanism adaptively suppresses modality noise by leveraging the semantic anchors as stable references. The framework further integrates contrastive learning to strengthen cross-modal alignment and employs a retrieve-then-read pipeline for open-domain answer reasoning. Experiments on the OK-VQA, FVQA, and A-OKVQA datasets demonstrate that the proposed framework outperforms state-of-the-art methods across multiple metrics, validating the effectiveness and robustness of the proposed framework.
JunMing et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: