PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 24, 20241 citationsOpen Access

Fusion of Domain-Adapted Vision and Language Models for Medical Visual Question Answering

View Full Paper
CHCuong HaSAShima AsaadiSKSanjeev Kumar Karn

Key Points

Key points are not available for this paper at this time.

Abstract

Vision-language models, while effective in general domains and showing strong performance in diverse multi-modal applications like visual question-answering (VQA), struggle to maintain the same level of effectiveness in more specialized domains, e.g., medical. We propose a medical vision-language model that integrates large vision and language models adapted for the medical domain. This model goes through three stages of parameter-efficient training using three separate biomedical and radiology multi-modal visual and text datasets. The proposed model achieves state-of-the-art performance on the SLAKE 1.0 medical VQA (MedVQA) dataset with an overall accuracy of 87.5% and demonstrates strong performance on another MedVQA dataset, VQA-RAD, achieving an overall accuracy of 73.2%.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ha et al. (2024) studied this question.

synapsesocial.com/papers/68e6de67b6db64358765a020https://doi.org/10.48550/arxiv.2404.16192
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Beyond the Hype: A dispassionate look at vision-language models in medical scenario2024
  2. 2Systematic Analysis of Vision–Language Models for Medical Visual Question Answering2026 · 1 citations
  3. 3Medical Vision-Language Modeling With Semantic Interaction and Adaptive Refinement Prompting for Bias Mitigation2025
  4. 4A Vision-Language Model with Multi-Granular Knowledge Fusion in Medical Imaging2024
  5. 5Medical Vision-Language Models: Existing Technologies, Clinical Applications and Future Directions2026