PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 18, 2026Scottish Medical Journal0 citations

Evaluation of large language models with clinical guidance for vetting outpatient magnetic resonance imaging lumbar spine referrals

View Full Paper
WCWilliam ClackettNHS TaysideHAHatim AlsusaNHS TaysideHWHannah WatsonCanon (France)

Key Points

  • The aim is to assess how accurately and quickly large language models can vet lumbar spine MRI referrals for sciatica.
  • Three LLMs (GPT-4, Claude Opus, Gemini) were tested on MRI referrals.
  • Models assigned outcomes (Accept - Routine, Accept - Urgent, Reject) and flagged contraindications.
  • Radiology registrars synthesized 120 referrals vetted by board-certified radiologists.
  • Performance metrics included accuracy, precision, recall, and F1 scores.
  • Inter-rater agreement between radiologists was substantial (κ = 0.76 for outcomes, κ = 0.68 for contraindication detection).
  • Claude Opus achieved the highest accuracy of 0.86 in vetting outcomes.
  • GPT-4 had the highest F1 score of 0.88 for contraindication detection.
  • LLMs completed the vetting task significantly faster than radiologists (average 9.8 min vs 135.0 min).

Abstract

ObjectivesAccurate triage of lumbar spine magnetic resonance imaging (MRI) referrals for sciatica is important for patient assessment, diagnosis and surgical planning. This study evaluates the accuracy and speed of large language models (LLMs) in automatically vetting lumbar spine MRI referrals from general practice.MethodsThree LLMs (GPT-4, Claude Opus, Gemini) were tasked with assigning an outcome (Accept - Routine, Accept - Urgent, Reject) and flagging MRI contraindications for lumbar spine referrals. Three prompts of increasing detail, including clinical guidelines and training examples, were used. Two radiology registrars synthesised 120 referrals, vetted by two board-certified radiologists, with a third resolving disagreements. Performance was assessed using accuracy, precision, recall and F1 scores.ResultsInter-rater agreement between radiologists was substantial for vetting outcome (Cohen's κ = 0.76) and contraindication detection (κ = 0.68). Claude Opus with the full prompt achieved the highest accuracy (0.86) for vetting outcomes. GPT-4 with the instruction-only prompt achieved the highest F1 score (0.88) for contraindication detection. LLMs completed the task substantially faster than radiologists (9.8 ± 1.0 vs 135.0 ± 45.0 min).ConclusionsLLMs demonstrate promising performance in vetting radiological referrals for sciatica, particularly with detailed context. All models identified all urgent referrals, suggesting potential for prioritising vetting worklists and improving timeliness of care.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Clackett et al. (2026) studied this question.

synapsesocial.com/papers/69e3205140886becb653f68ehttps://doi.org/10.1177/00369330261441582
Ask AI
Helpful
Bookmark
Share
View Full Paper