PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 3, 20260 citations

Measuring the gap: correlating synthetic-to-real drift with PHI de-identification performance.

View Full Paper
JCJoseph CorneliusFRFabio Rinaldi

Key Points

  • This research aims to assess how the drift between synthetic and real clinical notes affects de-identification performance for protected health information.
  • Generated synthetic clinical notes across four categories using five generator LLMs and one judge LLM.
  • Fine-tuned de-identification models on real, synthetic, and mixed corpora.
  • Evaluated model performance on three external benchmarks under a harmonized label schema.
  • Synthetic data shows moderate correlation with external benchmarks, specifically an out-of-distribution F1 score.
  • Models trained on broad clinical sources perform better than those using legal or narrowly synthetic data.
  • Drift measures can provide insights into improving synthetic data quality control.

Abstract

Clinical text de-identification enables the use of electronic health records while protecting patient privacy, but public training data remain scarce and often have mismatched documentation styles. Recent works have proposed using large language models (LLMs) to generate synthetic clinical notes, but it remains unclear if they reflect distributions of real clinical notes. We examine how lexical and semantic drift across training and evaluation corpora affects protected health information (PHI) tagger performance. We generated synthetic notes from scratch for four categories using five generator LLMs and one judge LLM. Next, we fine-tuned small de-identification models on real, synthetic, and mixed corpora, and evaluated them on three external benchmarks under a harmonized label schema. Models trained on broad, clinically oriented sources transfer better than those on legal or narrowly synthetic data. These results suggest that although synthetic data lacks some real-world distributional properties, it remains useful in low-resource settings. We found that compact distributional and embedding-based drift measures moderately correlate with out-of-distribution F1 score, a practically important result because drift estimation can improve synthetic-data quality control and alignment.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Cornelius et al. (2026) studied this question.

synapsesocial.com/papers/69f6e5618071d4f1bdfc60echttps://doi.org/10.1186/s44342-026-00072-9
Ask AI
Helpful
Bookmark
Share
View Full Paper