PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
December 2, 2025Nature Communications3 citationsOpen Access

A Benchmark for Breast Cancer Screening and Diagnosis in Mammogram Visual Question Answering

View Full Paper
JZJiayi ZhuFHFuxiang HuangQLQiong Luo

Key Points

  • Large vision-language models show diagnostic performance equivalent to random guessing, indicating unreliability.
  • MammoVQA dataset unifies 15 public datasets, including 131,847 images for mammogram visual question answering.
  • LLaVA-Mammo achieves +19.66% weighted accuracy gains in internal validation, outperforming previous models.
  • Emphasizes the importance of standardized evaluation benchmarks for mammogram interpretation.

Abstract

Breast cancer remains the most prevalent malignancy in women worldwide. Mammography-based early detection plays a pivotal role in improving patient survival outcomes. While large vision-language models offer transformative potential for mammogram visual question answering, the absence of standardized evaluation benchmarks currently makes it hard to fairly compare different large vision-language models' performance in mammogram interpretation. In this study, we address this critical gap through three key contributions: (1) We introduce MammoVQA, a mammogram visual question-answering dataset that unifies 15 public datasets, comprising 131,847 images (421K question-answering pairs) for image-level cases and 72,518 exams (476K images, 144K question-answering pairs) for exam-level cases. (2) Systematic evaluation of 12 recent high-performance large vision-language models (6 general, 6 medical) reveals diagnostic performance statistically equivalent to random guessing, highlighting their unreliability for mammogram interpretation. (3) Our domain-optimized LLaVA-Mammo achieves average +19.66% weighted accuracy gains over the best recent high-performance model in internal validation, with average +21.21% weighted accuracy improvements in external validation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhu et al. (2025) studied this question.

synapsesocial.com/papers/692e3d626c9b3ab28c186cd4https://doi.org/10.1038/s41467-025-66507-z
Ask AI
Helpful
Bookmark
Share
View Full Paper