PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 28, 2026JMIR Formative Research0 citationsOpen Access

Evaluation of the Accuracy of Probabilistic Record Linkage Across Sociodemographic Categories in 4 Databases: Exploratory Study

View Full Paper
CBCristina BarboiFOFangqian OuyangLLLauren R. Lembcke

Key Points

  • This study explores the accuracy of probabilistic patient matching across demographic categories, focusing on age, sex, race, and ethnicity.
  • Utilized 4 Indiana data sources, including health and death records.
  • Applied a modified Fellegi-Sunter probabilistic linkage algorithm for matching.
  • Established gold standard match status through dual manual review and adjudication.
  • Estimated matching performance metrics (sensitivity, positive predictive value, F1-scores) across demographic strata.
  • The matching F1-scores exceeded 0.82 for all age groups, with varied performance by sex, race, and ethnicity.
  • Sensitivity ranged from 0.70 to 0.97, showing discrepancies based on missing data.
  • Greater missingness correlated with lower matching accuracy, particularly noted in Newborn Screening and Death Master File.
  • Race and ethnicity data had the highest missingness and least informational diversity, affecting accuracy.

Abstract

Abstract Background Accurate patient record linkage is essential for clinical care, health information exchange, research, and public health surveillance. However, linkage accuracy may vary across demographic groups due to differences in data completeness, quality, and the structural factors underlying how demographic information is captured. Objective This study aimed to explore whether probabilistic patient matching accuracy varies by age, sex, race, and ethnicity and to identify potential sources of bias that may influence matching performance. Methods We used 4 Indiana data sources—the Indiana Network for Patient Care, Newborn Screening, Social Security Administration Death Master File, and Marion County Public Health Department—and applied a modified Fellegi-Sunter probabilistic linkage algorithm accommodating missing data under a missing at random assumption. Gold standard match status was established through dual manual review with adjudication. For each dataset, matching sensitivity, positive predictive value, and F 1 -scores were estimated and stratified by age, sex, race, and ethnicity. Data completeness, distinct value ratio, and Shannon entropy were assessed to characterize data quality. Ninety-five percent bootstrap CIs were used to assess significance. Results The algorithm-matching F 1 -score was greater than 0.82 for all age strata, ranging from 0.88 to 0.97 for sex, 0.85 to 0.99 for race, and 0.88 to 0.99 for ethnicity. Sensitivity ranged from 0.70 to 0.97 across age strata, 0.76 to 0.97 across sex, 0.85 to 0.99 across race, and 0.85 to 0.989 across ethnicity. Lower sensitivity and F 1 -scores were consistently observed in strata with greater missingness or discordance, particularly in Newborn Screening and Social Security Administration Death Master File. Race and ethnicity exhibited the highest missingness and lowest informational diversity, coinciding with the largest declines in accuracy. Shannon entropy and distinct value ratio varied across demographic groups and were strongly associated with performance, indicating that both low and excessively high informational diversity can impair matching. Conclusions Probabilistic patient matching accuracy is not uniform across demographics and is strongly influenced by data quality and completeness. Although overall matching performance, as assessed by the F 1 -score, remained above 0.8, it varied across datasets when stratified by sociodemographic characteristics. Sociodemographic data missingness is associated with lower matching accuracy, raising equity and ethical concerns for clinical, research, and public health applications. Routine demographic-stratified evaluations of matching accuracy, improved standardization of sociodemographic data, and fairness-aware linkage methods are essential to prevent the amplification of structural inequities in linked health datasets.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Barboi et al. (2026) studied this question.

synapsesocial.com/papers/69a286da0a974eb0d3c0227fhttps://doi.org/10.2196/78622
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Dissecting racial bias in an algorithm used to manage the health of populations2019 · 6,924 citations
  2. 2Race/ethnicity and the 2000 census: implications for public health2000 · 78 citations
  3. 3Mining for equitable health: Assessing the impact of missing data in electronic health records2023 · 126 citations
  4. 4Algorithmic fairness in computational medicine2022 · 159 citations
  5. 5Data Quality in Electronic Health Record Research: An Approach for Validation and Quantitative Bias Analysis for Imperfectly Ascertained Health Outcomes Via Diagnostic Codes2022 · 22 citations