Vision-Language Models (VLMs) have demonstrated impressive capabilities across various medical tasks, including report generation and visual question answering (VQA). However, pixel-level tasks such as image segmentation remain relatively underexplored, despite their critical importance for clinical decision-making, surgical planning, and model interpretability. Moreover, the scarcity of high-quality segmentation annotations in the medical domain often leads to biased data distributions, characterized by imbalances in disease types, anatomical coverage, and image quality. These biases are frequently overlooked during both model development and evaluation, limiting the robustness and real-world applicability of VLMs in healthcare scenarios. In this study, we propose a unified medical vision-language model applicable for a variety of clinical tasks, including report generation, VQA, and pixel-level image segmentation. Within the model, we propose a semantic interaction mechanism aimed at enhancing pixel-level vision and language representation learning. To mitigate the impact of biased data distributions, we explicitly develop an adaptive refinement prompting method involving the iterative re-prompting of hard samples. The proposed method is thoroughly validated through experiments on eight datasets and comparisons with nine state-of-the-art methods. The experimental results indicate that our model achieves superior performance in both medical VQA and segmentation tasks. These results highlight the potential of our approach in advancing the deployment of medical VLMs in real-world clinical applications. Code will be released at: https://github.com/SZUHvern/Unified-Medical-Vision-Language-Modeling.
Li et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: