Comparison of GPT-4o and specialty-tuned o1 shows similar accuracy in guideline adherence for otolaryngology.
Artificial-intelligence chatbots are gaining prominence in otolaryngology, yet their clinical safety depends on strict adherence to practice guidelines. The authors compared the accuracy of OpenAI's general-purpose GPT-4o model with the specialty-tuned o1 model on 100 otolaryngology questions drawn from national guidelines and common clinical scenarios spanning 7 subspecialty domains. Blinded otolaryngologists graded each answer as correct, partially correct, incorrect, or non-answer (scores 1, 0.5, 0, respectively), and paired statistical tests assessed performance differences. The o1 model delivered fully correct responses for 73% of questions, partially correct for 26%, and incorrect for 1%, yielding a mean accuracy score of 0.86. GPT-4o produced 64% correct and 36% partially correct answers with no incorrect responses, for a mean score of 0.82. The 4-point gap was not statistically significant (paired t test P=0.165; Wilcoxon P=0.157). Pediatric questions had the highest correctness (o1=92.9%, GPT-4o=78.6%). No domain showed systematic critical errors. Both models thus supplied predominantly guideline-concordant information, and specialty tuning conferred only a modest, nonsignificant benefit in this data set. These findings suggest contemporary large-language models may approach reliability thresholds suitable for supervised decision support in otolaryngology, but continual validation and oversight remain essential before routine deployment.
No takes yet. Share an insight, caveat, or question.
Prasad et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: