PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 21, 2026Computer Methods and Programs in Biomedicine Update0 citationsOpen Access

Aligning language models with clinical expertise: Direct preference optimization for heart failure nursing documentation in critical care

View Full Paper
JFJunyi FanLSLi SunNANegin Ashrafi

Key Points

  • The aim is to enhance the quality of nursing documentation for heart failure management in ICUs using direct preference optimization techniques.
  • Applied Direct Preference Optimization to Mistral-7B language model.
  • Utilized 8838 heart failure nursing notes from the MIMIC-III database.
  • Generated 21,210 preference pairs from expert-verified outputs and original notes.
  • Evaluated improvements using BLEU, ROUGE, BERTScore, and qualitative assessments.
  • Established a novel error taxonomy for clinical documentation.
  • BLEU score increased by 84%, from 0.173 to 0.318.
  • BERTScore improved by 7.6%, from 0.828 to 0.891.
  • Expert ratings showed significant increases in accuracy (+14.4 points), completeness (+14.5 points), logical consistency (+14.1 points), readability (+11.1 points), and structural clarity (+6.0 points).
  • DPO reduced documentation omissions, contradictions, and structural inconsistencies.

Abstract

Nursing documentation in intensive care units (ICUs) provides essential clinical intelligence but often suffers from inconsistent terminology, informal styles, and lack of standardization, challenges that are particularly critical in heart failure care. This study applies Direct Preference Optimization (DPO) to adapt Mistral-7B, a locally deployable language model, using 8838 heart failure nursing notes from the MIMIC-III database and 21,210 preference pairs derived from expert-verified GPT outputs, model generations, and original notes. Evaluation across BLEU, ROUGE, BERTScore, Perplexity, and expert qualitative assessments demonstrates that DPO markedly enhances documentation quality. Specifically, BLEU increased by 84% (0.173 → 0.318), BERTScore improved by 7.6% (0.828 → 0.891), and expert ratings rose across accuracy (+14.4 points), completeness (+14.5 points), logical consistency (+14.1 points), readability (+11.1 points), and structural clarity (+6.0 points). These results indicate that DPO can align lightweight clinical language models with expert standards, supporting privacy-preserving, AI-assisted documentation within electronic health record systems to reduce administrative burden and improve ICU patient safety. • Present the first application of direct preference optimization to align a locally deployable 7B language model with expert-verified ICU heart failure nursing documentation standards. • Construct 21,210 structured preference pairs from expert-checked GPT outputs, original nursing notes, and model generations using 8838 MIMIC-III samples. • Demonstrate measurable improvements in BLEU, ROUGE, BERTScore, and perplexity compared with the untuned Mistral-7B baseline. • Show through blinded clinical assessment that DPO reduces omissions, contradictions, and structural inconsistencies while improving accuracy and completeness. • Provide a novel error taxonomy and qualitative analysis showing how DPO reshapes narrative structure and clinical clarity in nurse documentation.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Fan et al. (2026) studied this question.

synapsesocial.com/papers/69be34d16e48c4981c672fa0https://doi.org/10.1016/j.cmpbup.2026.100244
Ask AI
Helpful
Bookmark
Share
View Full Paper