PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 15, 2025npj Digital Medicine16 citationsOpen Access

Leveraging large language models for the deidentification and temporal normalization of sensitive health information in electronic health records

View Full Paper
HDHong DaiTMTatheer Hussain MirCCChing-Tai Chen

Key Points

  • Fine-tuning large language models significantly improves deidentification performance, especially with lower-rank adaptations, compared to traditional methods.
  • Top teams in a recent competition achieved macro-F1 scores exceeding 0.8, indicating high effectiveness in temporal normalization and sensitive health information management.
  • Using 3,244 pathology reports, the study assessed the impact of various training strategies and highlighted the role of data augmentation in enhancing large language model capabilities.
  • Balancing performance and legal requirements is crucial for effective deidentification in healthcare, aiming for privacy and interpretability with sensitive health information.

Abstract

Secondary use of electronic health record notes enhances clinical outcomes and personalized medicine, but risks sensitive health information (SHI) exposure. Inconsistent time formats hinder interpretation, necessitating deidentification and temporal normalization. The SREDH/AI CUP 2023 competition explored large language models (LLMs) for these tasks using 3,244 pathology reports with surrogated SHIs and normalized dates. The competition drew 291 teams; the top teams achieved macro-F1 scores >0.8. Results were presented at the IW-DMRN workshop in 2024. Notably, 77.2% used LLMs, highlighting their growing role in healthcare. This study compares competition results with in-context learning and fine-tuned LLMs. Findings show that fine-tuning, especially with lower-rank adaptation, boosts performance but plateaus or degrades in models over 6 B parameters due to overfitting. Our findings highlight the value of data augmentation, training strategies, and hybrid approaches. Effective LLM-based deidentification requires balancing performance with legal and ethical demands, ensuring privacy and interpretability in regulated healthcare settings.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Dai et al. (2025) studied this question.

synapsesocial.com/papers/68a365560a429f797332b1d5https://doi.org/10.1038/s41746-025-01921-7
Ask AI
Helpful
Bookmark
Share
View Full Paper