PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 13, 2026SHILAP Revista de lepidopterología0 citationsOpen Access

Evaluating the sampling effect of propensity score matching for reducing selection bias in medical data

MRMinji RohSYSujin YumGJGihun Joo

Key Points

  • The central aim is to evaluate the effectiveness of propensity score matching in mitigating selection bias in medical datasets.
  • Evaluated propensity score matching alongside various undersampling and oversampling techniques.
  • Applied methods to three medical datasets with different degrees of demographic selection bias.
  • Trained and compared six classification models to assess the impact of resampling techniques on performance.
  • Propensity score matching significantly reduces selection bias as measured by standardized mean difference (SMD).
  • It maintains stable classification performance under moderate demographic imbalance.
  • Overall model reliability and generalization potential improve with the application of propensity score matching.

Abstract

Background In real-world medical data, selection bias can significantly impact the performance of machine learning models, potentially leading to distorted outcomes. However, research aimed at mitigating selection bias remains relatively limited. Methods In this study, we evaluate the effectiveness of Propensity Score Matching (PSM) in reducing selection bias and assessing its impact on classification performance in imbalanced medical data. Specifically, we apply PSM alongside five undersampling, three oversampling, and three hybrid sampling techniques to three medical datasets: rapidly progressive dementia prediction (ADNI, n = 628, events = 51), hypothyroidism prediction (UCI, n = 3,772, events = 3,481), and cardiovascular disease prediction (Kaggle, n = 253,680, events = 23,893), each exhibiting varying degrees of demographic selection bias. We train and compare six classification models to assess the impact of each resampling technique on model performance. The magnitude of selection bias is quantified using the standardized mean difference (SMD), while model performance is assessed using the Area Under the Receiver Operating Characteristic Curve (AUROC), the Area Under the Precision-Recall Curve (AUPRC), accuracy, precision, recall, F1-score, specificity, calibration curves, Brier score, and decision curve analysis. Results The results indicate that PSM reduces SMD within the dataset, maintains stable classification performance, and enhances the internal validity of the model under conditions of limited or moderate demographic imbalance. Conclusion These advantages suggest its potential for improving model reliability and facilitating better generalization to external datasets in real-world medical applications. However, in datasets with extreme selection bias or when overly restrictive matching is applied, PSM can degrade model performance, underscoring the importance of choosing strategies that account for dataset characteristics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Roh et al. (2026) studied this question.

synapsesocial.com/papers/698ebedd85a1ff6a93016253https://doi.org/10.3389/fpubh.2026.1747762
Ask AI
Helpful
Bookmark
Share
View Full Paper