PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 6, 20250 citationsOpen Access

Evaluating Large Language Models for Automatic Detection of In-Hospital Cardiac Arrest: Multi-Site Analysis of Clinical Notes

View Full Paper
UVUğurcan VurgunAKAarthi KaviyarasuSHSy Hwang

Key Points

  • GPT-4o achieved an F1-score of 0.90 and recall of 0.97 for identifying in-hospital cardiac arrests.
  • Four language models were evaluated across 2,674 clinical notes from five hospitals, with notable performance variations.
  • Site-specific documentation practices strongly influenced model performance, with inter-site agreement rates differing over 20%.
  • Identified challenges include medical terminology hallucinations and structural inconsistencies in model reasoning during detection.

Abstract

In-hospital cardiac arrest (IHCA) affects over 200,000 patients annually in the United States, yet its detection through manual chart review remains resource-intensive and often delayed. We evaluated the performance of four open-source large language models (LLMs) and GPT-4o in identifying IHCA cases from 2,674 clinical notes across five hospitals. While GPT-4o achieved the highest performance (F1-score: 0.90, recall: 0.97), several open-source models demonstrated comparable capabilities, suggesting their viability for clinical applications. Our systematic analysis of model outputs revealed that performance was strongly influenced by site-specific documentation practices, with inter-site agreement rates varying by over 20%. Through detailed error analysis, we identified key challenges including medical terminology hallucinations and structural inconsistencies in model reasoning. These findings establish a framework for implementing LLM-based IHCA detection systems while highlighting critical considerations for their clinical deployment.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Vurgun et al. (2025) studied this question.

synapsesocial.com/papers/689523d29f4f1c896c42a0d8https://doi.org/10.1101/2025.08.04.25331524
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A large language model–based generative natural language processing framework fine‐tuned on clinical notes accurately extracts headache frequency from electronic health records2024 · 46 citations
  2. 2Abstract 328: Identification of In-Hospital Cardiac Arrest Using Administrative Billing Codes2023 · 2 citations
  3. 3Two-Layer Retrieval-Augmented Generation Framework for Low-Resource Medical Question Answering Using Reddit Data: Proof-of-Concept Study2024 · 22 citations
  4. 4Extraction of sleep information from clinical notes of Alzheimer’s disease patients using natural language processing2024 · 23 citations
  5. 5Digital oximetry biomarkers for assessing respiratory function: standards of measurement, physiological interpretation, and clinical use2021 · 128 citations