PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 28, 20241 citationsOpen Access

STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical

View Full Paper
GSGuohao SunCQCan QinHFHuazhu Fu

Key Points

Key points are not available for this paper at this time.

Abstract

Large Vision-Language Models (LVLMs) have shown significant potential in assisting medical diagnosis by leveraging extensive biomedical datasets. However, the advancement of medical image understanding and reasoning critically depends on building high-quality visual instruction data, which is costly and labor-intensive to obtain, particularly in the medical domain. To mitigate this data-starving issue, we introduce Self-Training Large Language and Vision Assistant for Medical (STLLaVA-Med). The proposed method is designed to train a policy model (an LVLM) capable of auto-generating medical visual instruction data to improve data efficiency, guided through Direct Preference Optimization (DPO). Specifically, a more powerful and larger LVLM (e.g., GPT-4o) is involved as a biomedical expert to oversee the DPO fine-tuning process on the auto-generated data, encouraging the policy model to align efficiently with human preferences. We validate the efficacy and data efficiency of STLLaVA-Med across three major medical Visual Question Answering (VQA) benchmarks, demonstrating competitive zero-shot performance with the utilization of only 9% of the medical data.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sun et al. (2024) studied this question.

synapsesocial.com/papers/68e62e92b6db6435875c058dhttps://doi.org/10.48550/arxiv.2406.19973
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Advancing High Resolution Vision-Language Models in Biomedicine2024 · 1 citations
  2. 2Prompting Medical Large Vision-Language Models to Diagnose Pathologies by Visual Question Answering2024
  3. 3Dr-LLaVA: Visual Instruction Tuning with Symbolic Clinical Grounding2024
  4. 4Beyond the Hype: A dispassionate look at vision-language models in medical scenario2024
  5. 5Medical Large Vision Language Models with Multi-Image Visual Ability2025