PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 5, 2026npj Digital Medicine0 citationsOpen Access

Intersectional fairness in vision-language models for medical image disease classification

YZYupeng ZhangADAdam G. DunnUNUsman Naseem

Key Points

  • To address intersectional biases in medical AI using a training framework that standardizes diagnostic performance across diverse patient groups.
  • Developed Cross-Modal Alignment Consistency (CMAC-MMD) framework to enhance diagnostic certainty without using sensitive demographic data.
  • Evaluated using 10,015 skin lesion images with external validation on additional datasets for diverse attributes.
  • Stratified performance based on age, gender, and race.
  • Reduced overall intersectional missed diagnosis gap (ΔTPR) in dermatology from 0.50 to 0.26, improving AUC from 0.94 to 0.97.
  • For glaucoma screening, reduced ΔTPR from 0.41 to 0.31, with AUC improving to 0.72 from a baseline of 0.71.

Abstract

Medical artificial intelligence (AI) systems, particularly multimodal vision-language models (VLM), often exhibit intersectional biases, with models systematically less confident in diagnosing marginalised patient subgroups and producing higher rates of missed diagnoses. Current fairness interventions frequently fail to address these gaps or compromise overall diagnostic performance to achieve statistical parity. In this study, we developed Cross-Modal Alignment Consistency (CMAC-MMD), a training framework that standardises diagnostic certainty across intersectional patient subgroups without requiring sensitive demographic data during clinical inference. We evaluated this approach using 10,015 skin lesion images (HAM10000) with external validation on 12,000 images (BCN20000), and 10,000 fundus images for glaucoma detection (Harvard-FairVLMed), stratifying performance by intersectional age, gender, and race attributes. In the dermatology cohort, the proposed method reduced the overall intersectional missed diagnosis gap (difference in True Positive Rate, ΔTPR) from 0.50 to 0.26 while improving the Area Under the Curve (AUC) from 0.94 to 0.97 compared to standard training. Similarly, for glaucoma screening, the method reduced ΔTPR from 0.41 to 0.31, achieving a better AUC of 0.72 (vs. 0.71 baseline). This provides a methodological foundation toward clinical decision support systems that are both accurate and perform more equitably across diverse patient subgroups, without increasing privacy risks during inference.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhang et al. (2026) studied this question.

synapsesocial.com/papers/6a72e7c1226790f370656ebchttps://doi.org/10.1038/s41746-026-03030-5
Ask AI
Helpful
Bookmark
Share
View Full Paper