PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 3, 2026Frontiers in Digital Health0 citationsOpen Access

Assessment of frontier Large Language Models in sleep medicine

APAnshum PatelJacksonville CollegeHCHet ContractorJacksonville CollegeHHHayden HeningerJacksonville College

Key Points

  • The research aims to evaluate the diagnostic performance of two large language models in sleep medicine.
  • Two models, ChatGPT-5 and Grok-4, were tested on diagnostic reasoning using 79 clinical vignettes and 897 multiple-choice questions.
  • Performance was evaluated through accuracy scoring for final diagnosis and scoring methods for differential diagnosis.
  • Inter-model performance used the Mann-Whitney U test for statistical comparison.
  • Both models achieved 92.4% accuracy for final diagnosis (95% CI 86.4, 98.4) and high scores in MCQs (ChatGPT-5: 93.0%; Grok-4: 92.8%).
  • F1-scores for differential diagnosis were modest (ChatGPT-5: 0.55 ± 0.20; Grok-4: 0.59 ± 0.20).
  • There were no statistically significant differences in performance between the two models (p > 0.05).

Abstract

Study objectives To evaluate and compare the performance of two proprietary frontier large language models (LLMs), ChatGPT-5 and Grok-4, on diagnostic reasoning and foundational knowledge tasks within the specialty domain of sleep medicine. Methods The models were evaluated on two tasks: case-based reasoning using 79 clinical vignettes from the AASM Case Book of Sleep Medicine and knowledge assessment using 897 multiple-choice questions (MCQs) from board review materials. For vignettes, final diagnosis was scored by concept-level exact match, and differential diagnosis (DDx) was scored on a fixed top-5 output using concept-level matching with synonym normalization to compute precision, recall, and F 1-score. MCQ performance was the proportion correct. Inter-model performance was compared using the Mann–Whitney U test. Results Both models achieved high accuracy for final diagnosis (92.4% for both; 95% CI 86.4, 98.4) and MCQs (ChatGPT-5: 93.0%; Grok-4: 92.8%). However, performance on generating a comprehensive differential diagnosis was suboptimal, with modest F 1-scores for both ChatGPT-5 (0.55 ± 0.20) and Grok-4 (0.59 ± 0.20). There were no statistically significant differences in performance between the two models across any metric ( p 0.05). Conclusions Frontier LLMs demonstrated high accuracy in sleep medicine tasks requiring knowledge recall and direct pattern recognition but showed more limited performance in complex clinical reasoning tasks such as generating a comprehensive differential diagnosis. These findings suggest that current general-purpose models may be more reliable for focused knowledge support than for broad hypothesis generation. Future studies should evaluate whether domain-adapted models or clinician-in-the-loop workflows can improve real-world diagnostic usefulness and safety.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Patel et al. (2026) studied this question.

synapsesocial.com/papers/69f6e5308071d4f1bdfc5f4bhttps://doi.org/10.3389/fdgth.2026.1769386
Ask AI
Helpful
Bookmark
Share
View Full Paper