This is the authors' abstract. We don't add key points for this paper.
The advent of artificial intelligence has precipitated the gradual implementation of visual question answering (VQA) in the domain of medical imaging. However, extant models are deficient in terms of the specialisation and temporal reasoning capabilities that are prerequisites for the medical field. The proposed framework is a multimodal visual question answering framework based on medical knowledge graphs and large language models, termed MediLKG-VQA(Medical LLM-Knowledge-Grounded VQA). This framework integrates a medical knowledge graph, constructed utilising the MIMIC-CXR-JPG dataset, with a multirelationship graph neural network, with the objective of enhancing semantic reasoning capabilities. The model’s capacity to analyse variations in medical images across different time points is facilitated by dual-graph input, thereby enabling the inference of lesion progression. Furthermore, the incorporation of the Qwen-3 large language model facilitates bilingual (Chinese-English) question-answering and structured explanations, thereby enhancing the system’s interactivity and clinical adaptability. Experimental findings demonstrate that the system exhibits superior performance in terms of accuracy, interpretability, and clinical adaptability when compared with existing methods. This provides a novel solution for intelligent diagnosis and medical decision-making.
No takes yet. Share an insight, caveat, or question.
Wang et al. (2025) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: