PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 26, 2026Diagnostics1 citationsOpen Access

Performance of Multimodal Large Language Models in Detection and Position Assessment of Thoracic Devices on Chest Radiographs

View Full Paper
HGHamza Eren GüzelIzmir UniversityCÖCemre ÖzenbaşIzmir UniversityBSBabak SaraviDüsseldorf University Hospital

Key Points

  • To assess the performance of multimodal large language models in detecting and positioning thoracic devices on chest radiographs.
  • Evaluation of three LLMs on 4813 chest radiographs from the RANZCR CLiP dataset.
  • Quantified performance using balanced accuracy, MCC, and kappa statistics.
  • Conducted additional analyses including a blinded reader study and prompt-sensitivity analysis.
  • Device presence performance varied; abnormal-position sensitivity was poor (MCC ≤ 0.028; balanced accuracy 0.41–0.53).
  • Radiologists outperformed all LLMs in 42/42 comparisons, with 33 showing statistical significance (p < 0.001).
  • Inter-model agreement was poor (Fleiss’ κ: 0.005–0.383 for device presence).

Abstract

Background: Accurate identification and positioning of thoracic devices on chest radiographs is critical for patient safety in intensive care. Multimodal large language models (LLMs) offer potentially generalizable automated evaluation, but their performance in this domain is underexplored. Methods: Three multimodal LLMs (GPT-4o, gpt-4o-2024-08-06; Gemini 3.1 Flash Lite Preview; Claude Sonnet 4.6) were evaluated on 4813 chest radiographs from the RANZCR CLiP dataset for device presence and positioning of ETT, NGT, CVC, and Swan–Ganz catheters. Performance was quantified with 95% Wilson confidence intervals, balanced accuracy, MCC, Cochran’s Q, Bonferroni-corrected McNemar, and Cohen’s/Fleiss’ kappa. Six additional analyses were performed: a blinded paired reader study (n = 377; two board-certified radiologists, blinded to ground truth and to all LLM outputs), external validation on PadChest (n = 200, device-presence detection only—PadChest lacks granular position labels), three-variant prompt-sensitivity analysis (n = 103), repeat-inference stability across three runs (n = 50), systematic error taxonomy, and a failure-case analysis. Results: Device-presence performance varied widely across models; abnormal-position sensitivity was uniformly poor (MCC ≤ 0.028; balanced accuracy 0.41–0.53). Inter-model agreement was poor to slight (Fleiss’ κ: 0.005–0.383 for presence; −0.280 to −0.025 for classification). Radiologists numerically outperformed all three LLMs in 42/42 paired comparisons; the superiority was statistically significant after Bonferroni correction in 33/42 (32/42 at p < 0.001). PadChest replicated the negative finding for device-presence detection (malposition not externally validated). Prompts and inference stochasticity introduced 2–3× sensitivity swings and run-to-run κ from 0.20 to 0.85. Case failures concentrated systematically in multi-device cases (p < 0.0001) but not in abnormal-position cases (p = 0.14). Conclusions: Current general-purpose multimodal LLMs are not yet reliable for autonomous thoracic-device assessment; their failure patterns are structurally characterizable across models, prompts, and case types and support, at most a circumscribed role, as adjunct device-presence screening tools. The findings do not generalize to purpose-built, regulator-approved clinical AI systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Güzel et al. (2026) studied this question.

synapsesocial.com/papers/6a153b00b5d9c58d83e8d425https://doi.org/10.3390/diagnostics16111602
Ask AI
Helpful
Bookmark
Share
View Full Paper