What question did this study set out to answer?

The aim is to enhance generalization and reliability of medical reasoning using a novel reinforcement learning approach.

February 6, 2026

Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models

Key Points

The aim is to enhance generalization and reliability of medical reasoning using a novel reinforcement learning approach.
Development of Med-R1, a reinforcement learning-enhanced vision-language model.
Utilization of Group Relative Policy Optimization for reward-guided learning.
Evaluation across eight medical imaging modalities for comprehensive assessment.
Assessment of cross-task generalization by testing on five question types.
Comparison of Med-R1's performance against baseline models with different parameter sizes.
Med-R1 achieves a 29.94% improvement in average accuracy over the baseline Qwen2-VL-2B.
Outperforms the larger Qwen2-VL-72B model in several tasks.
Demonstrates a 32.06% improvement in question-type generalization over Qwen2-VL-2B.
Findings indicate that reasoning quality and position influence performance, challenging existing assumptions.

Abstract

Vision-language models (VLMs) have achieved impressive progress in natural image reasoning, yet their potential in medical imaging remains underexplored. Medical vision-language tasks demand precise understanding and clinically coherent answers, which are difficult to achieve due to complexity of medical data and the scarcity of high-quality expert annotations. These challenges limit the effectiveness of conventional supervised fine-tuning (SFT) and Chain-of-Thought (CoT) strategies that work well in general domains. To address these challenges, we propose Med-R1, a reinforcement learning (RL)-enhanced VLM designed to improve generalization and reliability in medical reasoning. Med-R1 adopts Group Relative Policy Optimization (GRPO) to encourage reward-guided learning beyond static annotations. We comprehensively evaluate Med-R1 across eight distinct medical imaging modalities. Med-R1 achieves a 29.94% improvement in average accuracy over its base model Qwen2-VL-2B, and even outperforms Qwen2-VL-72B-a model with 36× more parameters. To assess cross-task generalization, we further evaluate Med-R1 on five question types. Med-R1 outperforms Qwen2-VL-2B by 32.06% in question-type generalization, also surpassing Qwen2-VL-72B. We further explore the thinking process in Med-R1, a crucial component of Deepseek-R1. Our results show that omitting intermediate rationales (No-Thinking Med-R1) not only improves cross-domain generalization with less training, but also challenges the common assumption that more reasoning always helps. Nevertheless, we also find that the Think-After Med-R1 variant further improves performance while maintaining interpretability. These findings suggest that, in medical VQA, the mere presence of explicit reasoning does not guarantee better performance. Instead, performance depends on the quality of the reasoning and the position where the reasoning is generated.

Perguntar à IA

Bookmark