PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 7, 2026Diagnostics0 citationsOpen Access

Leveraging Large Language Models for Automated Extraction of Abdominal Aortic Aneurysm Features from Radiology Reports

View Full Paper
PMPraneel MukherjeeAlbert Einstein College of MedicineRLR. LeeAlbert Einstein College of MedicineRHRoham HadidchiAlbert Einstein College of Medicine

Key Points

  • This research aims to assess the effectiveness of large language models in extracting relevant information about abdominal aortic aneurysms from CT radiology reports.
  • Retrospective analysis of 500 abdominal CT reports mentioning AAA from 2020 to 2024.
  • Ground truth labels were established through manual review of reports.
  • Four open-source large language models were tested for extraction tasks.
  • Model outputs were assessed for exact-match accuracy and inter-model agreement using Fleiss’ kappa.
  • Accuracy for identifying AAA presence ranged from 0.90 to 0.95.
  • Accuracy for prior aortic repair ranged from 0.90 to 0.97.
  • Aneurysm size extraction showed accuracy between 0.67 and 0.88, influenced by class imbalance.
  • Rupture and impending rupture identification exceeded 0.90 accuracy across models, with lower agreement.

Abstract

Background/Objectives. Abdominal computed tomography (CT) radiology reports contain critical information for abdominal aortic aneurysm (AAA) management, including aneurysm presence, size, rupture status, and prior repair. However, this information is often embedded within lengthy, heterogeneous reports, making manual extraction inefficient. We evaluated the performance of multiple large language models (LLMs) for automated extraction of AAA-related findings from radiology reports. Methods. We retrospectively analyzed 500 abdominal CT reports mentioning AAA from an urban academic health system (2020–2024). Ground truth labels were established by manual review. Four open-source LLMs (Qwen2.5-7B-Instruct, Llama3-Med42-8B, GPT-OSS-20B, and MedGemma-27B-text-it) were evaluated for extraction of aneurysm presence, size, morphology, rupture status, impending rupture, and prior aortic repair. Model outputs were compared with ground truth using exact-match accuracy, and inter-model agreement was assessed using Fleiss’ kappa. Reasoning traces were examined to characterize correct and incorrect model behavior. Results. Accuracy for identifying AAA presence ranged from 0.90 to 0.95 (κ = 0.851), and prior aortic repair from 0.90 to 0.97 (κ = 0.793). Accuracy for aneurysm size ranged from 0.67 to 0.88 (κ = 0.340), with low κ’s due to class imbalance or dimension misselection. Rupture and impending rupture were identified with accuracies exceeding 0.90 across models, though agreement was lower (κ = 0.485 and 0.589), reflecting low event prevalence. Larger models (GPT-OSS-20B, MedGemma-27B) generally outperformed smaller models. Reasoning analysis revealed strengths in measurement prioritization but recurrent errors, including dimension misselection, over-inference of prior repair, and conservative classification of rupture-related findings. Conclusions. LLMs can accurately extract clinically relevant AAA information from radiology reports with interpretable reasoning, with larger and medically trained models outperforming smaller or general-purpose models. Performance varies by task and model, underscoring the need for careful validation and human-in-the-loop deployment in clinical settings.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mukherjee et al. (2026) studied this question.

synapsesocial.com/papers/69d49f8ab33cc4c35a227f96https://doi.org/10.3390/diagnostics16071083
Ask AI
Helpful
Bookmark
Share
View Full Paper