PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 5, 2025IEEE Journal of Biomedical and Health Informatics22 citationsOpen Access

Med3DVLM: An Efficient Vision-Language Model for 3D Medical Image Analysis

View Full Paper
XYXin YuGAGorkem Can AtesKGKuang Gong

Key Points

  • Med3DVLM combines efficient encoding and innovative learning strategies to enhance 3D image analysis.
  • It outperforms the previous state-of-the-art M3D-LaMed in tasks like image-text retrieval with 61.00% R@1.
  • Evaluated on the large M3D dataset, it demonstrates significant advances in report generation and VQA.
  • The model enables scalable reasoning across diverse clinical applications, supporting better integration of imaging and language.

Abstract

Vision-language models (VLMs) have shown promise in 2D medical image analysis, but extending them to 3D remains challenging due to the high computational demands of volumetric data and the difficulty of aligning 3D spatial features with clinical text. We present Med3DVLM, a 3D VLM designed to address these challenges through three key innovations: (1) DCFormer, an efficient encoder that uses decomposed 3D convolutions to capture fine-grained spatial features at scale; (2) SigLIP, a contrastive learning strategy with pairwise sigmoid loss that improves image-text alignment without relying on large negative batches; and (3) a dual-stream MLP-Mixer projector that fuses low- and high-level image features with text embeddings for richer multi-modal representations. We evaluated our model on the M3D dataset, which includes radiology reports and VQA data for 120,084 3D medical images. The results show that Med3DVLM achieves superior performance on multiple benchmarks. For image-text retrieval, it reaches 61.00% R@1 on 2,000 samples, significantly outperforming the current state-of-the-art M3D-LaMed model (19.10%). For report generation, it achieves a METEOR score of 36.42% (vs. 14.38%). In open-ended visual question answering (VQA), it scores 36.76% METEOR (vs. 33.58%), and in closed-ended VQA, it achieves 79.95% accuracy (vs. 75.78%). These results demonstrate Med3DVLM's ability to bridge the gap between 3D imaging and language, enabling scalable, multi-task reasoning across clinical applications. Our code is publicly available at https://github.com/mirthAI/Med3DVLM.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yu et al. (2025) studied this question.

synapsesocial.com/papers/68bb4d206d6d5674bcd00ea5https://doi.org/10.1109/jbhi.2025.3604595
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Towards Generalist Biomedical AI2024 · 403 citations
  2. 2Towards generalist foundation model for radiology by leveraging web-scale 2D&3D medical data2025 · 93 citations
  3. 3MedCLIP: Contrastive Learning from Unpaired Medical Images and Text2022 · 702 citations
  4. 4Sigmoid Loss for Language Image Pre-Training2023 · 879 citations
  5. 5VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge2025 · 20 citations