PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 10, 2026Health Information Science and Systems1 citationsOpen Access

Leveraging electronic health records for atrial fibrillation cohort generation

ADAne García Domingo-AldamaMPMarcos Merino PradoAGAlain García-Olea

Key Result

An adapted rule-based approach for cohort selection in Atrial Fibrillation achieved 82% accuracy, while LLMs reached 79% with significantly less manual effort involved.

Key Points

  • This study aims to enhance cohort selection for clinical trials by using NLP and LLMs, particularly in the context of atrial fibrillation and heart failure.
  • Annotated a dataset of 212 patients using discharge reports for AF progression.
  • Evaluated an adapted rule-based approach and zero-shot LLMs with varied prompts.
  • Annotated an additional dataset of 100 patients for HF decompensation to assess generalizability.
  • The rule-based approach achieved the highest accuracy of 0.82.
  • LLMs with prompt language strategies performed comparably with an accuracy up to 0.79.
  • The medium-sized general-domain model gemma-3 outperformed other models.

Structured PICO

Can Large Language Models and rule-based NLP effectively automate cohort selection for Atrial Fibrillation progression and Heart Failure decompensation from non-English electronic health records?

P
Population
212 patients with documented Atrial Fibrillation (AF) onset and 100 patients with documented Heart Failure (HF) first episode from the Basque Public Healthcare System (Osakidetza) electronic health records.
I
Intervention
Large Language Models (LLMs) including gemma-3, Llama-3.1, and Mistral with varying prompt structures (concatenation, summarization, onset-guided) for automated cohort selection.
C
Comparator
Adapted rule-based Natural Language Processing (NLP) pipeline.
O
Outcome
Accuracy and F1-score for identifying AF progression and HF decompensation from discharge reports.

Discharge reports combined with rule-based NLP or LLMs using task-division prompts can effectively automate complex, temporally-dependent cohort selection for conditions like AF progression and HF decompensation.

Limitations

  • LLMs struggled with long-context inputs
  • Prompt language strongly influenced performance
  • Medical model variants were not consistently superior

Abstract

Abstract Purpose Cohort selection and eligibility screening are critical in clinical research, especially in trials where manual patient matching remains a major bottleneck. This study investigates the use of Natural Language Processing and Large Language Models (LLMs) in two real use cases, namely Atrial Fibrillation (AF) progression and Hearth Failure (HF) decompensation, within a non-English clinical context. We specifically address the following research questions: (1) Can discharge reports and NLP support cohort selection? (2) Can LLMs effectively model longitudinal patient trajectories and temporal reasoning? (3) Do general-purpose or domain-adapted LLMs outperform rule-based baselines for this task? (4) Compared to large foundation models, do small-scale LLMs offer similar performance? Methods A dataset of 212 patients was manually annotated for AF progression using discharge reports. Two strategies were evaluated: (1) an adapted rule-based pipeline and (2) zero-shot open-source LLMs with varying prompt structures. To assess generalizability, an additional dataset of 100 patients was annotated for HF decompensation. Results The adapted rule-based approach achieved the highest accuracy (0.82), but LLMs with task-division prompts performed comparably (up to 0.79), requiring significantly less manual effort. The medium-sized general-domain gemma-3 model outperformed others. Conclusions (1) Discharge reports are a valuable resource for automatic cohort selection, with both the rule-based method and LLMs showing promising results. (2) While LLMs struggled with long-context inputs, they handled temporal reasoning well when explicit dates were provided. (3) Larger models did not always outperform smaller ones, (4) prompt language strongly influenced performance, and medical model variants were not consistently superior.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Domingo-Aldama et al. (2026) studied this question. An adapted rule-based approach for cohort selection in Atrial Fibrillation achieved 82% accuracy, while LLMs reached 79% with significantly less manual effort involved.

synapsesocial.com/papers/696321c391e05aa366cb8050https://doi.org/10.1007/s13755-025-00415-w
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Identification of recurrent atrial fibrillation using natural language processing applied to electronic health records2023 · 22 citations
  2. 2An Innovative Approach to Using Electronic Health Records Through Health Information Exchange to Build a Chronic Disease Registry in Michigan2024 · 6 citations
  3. 3Engineering of Generative Artificial Intelligence and Natural Language Processing Models to Accurately Identify Arrhythmia Recurrence2024 · 7 citations
  4. 4Maximizing clinical cohort size using free text queries2015 · 8 citations
  5. 5Automated disease cohort selection using word embeddings from Electronic Health Records2017 · 72 citations