PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
November 30, 2025Scientific Reports3 citationsOpen Access

Improving social determinants of health documentation in French electronic health records using large language models

View Full Paper
ABAdrien BazogePBPacôme Constant Dit BeaufilsMHMohammed Hmitouch

Key Points

  • The model achieved identification of social determinants of health in 95.8% of patients, compared to only 2.8% using traditional coding.
  • Performance was strong for well-documented categories like marital status and alcohol use with an F1 score above 0.80.
  • Using the Flan-T5-Large model, we trained on clinical notes from Nantes University Hospital to extract SDoH data.
  • These findings highlight the role of natural language processing in enhancing real-world health data in non-English EHR systems.

Abstract

Abstract Social determinants of health (SDoH) significantly influence health outcomes, shaping disease progression, treatment adherence, and health disparities. However, their documentation in structured electronic health records (EHRs) is often incomplete or missing. This study presents an approach based on large language models (LLMs) for extracting 13 SDoH categories from French clinical notes. We trained Flan-T5-Large on annotated social history sections from clinical notes at Nantes University Hospital, France. We evaluated the model at two levels: (i) identification of SDoH categories and associated values, and (ii) extraction of detailed SDoH with associated temporal and quantitative information. The model performance was assessed across four datasets, including two that we publicly release as open resources. The model achieved strong performance for identifying well-documented categories such as living condition, marital status, descendants, job, tobacco, and alcohol use (F1 score > 0.80). Performance was lower for categories with limited training data or highly variable expressions, such as employment status, housing, physical activity, income, and education. Our model identified 95.8% of patients with at least one SDoH, compared to 2.8% for ICD-10 codes from structured EHR data. Our error analysis showed that performance limitations were linked to annotation inconsistencies, reliance on English-centric tokenizer, and reduced generalizability due to the model being trained on social history sections only. These results demonstrate the effectiveness of NLP in improving the completeness of real-world SDoH data in a non-English EHR system.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bazoge et al. (2025) studied this question.

synapsesocial.com/papers/692b944c1d383f2b2a378cbdhttps://doi.org/10.1038/s41598-025-29987-z
Ask AI
Helpful
Bookmark
Share
View Full Paper