PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 2, 202541 citationsOpen Access

MedGemma Technical Report

View Full Paper
ASAndrew SellergrenSKSahar KazemzadehTJTiam Jaroensri

Key Points

  • MedGemma significantly exceeds the performance of similar-sized generative models, improving medical reasoning and understanding.
  • The collection achieves up to 18.1% improvement on medical tasks like chest X-ray classification and multimodal question answering.
  • Fine-tuning MedGemma reduces errors in electronic health records by 50%, achieving comparable performance to specialized methods.
  • MedSigLIP enhances visual understanding, demonstrating strong capabilities alongside MedGemma for diverse medical applications.

Abstract

Artificial intelligence (AI) has significant potential in healthcare applications, but its training and deployment faces challenges due to healthcare's diverse data, complex tasks, and the need to preserve privacy. Foundation models that perform well on medical tasks and require less task-specific tuning data are critical to accelerate the development of healthcare AI applications. We introduce MedGemma, a collection of medical vision-language foundation models based on Gemma 3 4B and 27B. MedGemma demonstrates advanced medical understanding and reasoning on images and text, significantly exceeding the performance of similar-sized generative models and approaching the performance of task-specific models, while maintaining the general capabilities of the Gemma 3 base models. For out-of-distribution tasks, MedGemma achieves 2.6-10% improvement on medical multimodal question answering, 15.5-18.1% improvement on chest X-ray finding classification, and 10.8% improvement on agentic evaluations compared to the base models. Fine-tuning MedGemma further improves performance in subdomains, reducing errors in electronic health record information retrieval by 50% and reaching comparable performance to existing specialized state-of-the-art methods for pneumothorax classification and histopathology patch classification. We additionally introduce MedSigLIP, a medically-tuned vision encoder derived from SigLIP. MedSigLIP powers the visual understanding capabilities of MedGemma and as an encoder achieves comparable or better performance than specialized medical image encoders. Taken together, the MedGemma collection provides a strong foundation of medical image and text capabilities, with potential to significantly accelerate medical research and development of downstream applications. The MedGemma collection, including tutorials and model weights, can be found at https://goo.gle/medgemma.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sellergren et al. (2025) studied this question.

synapsesocial.com/papers/68de5d9383cbc991d0a2009ahttps://doi.org/10.48550/arxiv.2507.05201
Ask AI
Helpful
Bookmark
Share
View Full Paper