Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 30, 2026Applied Clinical Informatics

A Comparative Analysis of Large Language Model Performance on USMLE Step 1-Style Allergy/Immunology Questions: Evaluating Correctness and Consistency

View Full Paper
Ask AI
Bookmark
Share

Authors

MCM. CarrollSKSabrina KentisHKHannah Kareff

Discussion

Loading...

Member takes

Overview

Benchmarking study demonstrates variable accuracy and consistency across prompt formats on medical licensing questions, highlighting prompt sensitivity in AI education.

Key Points

  • To evaluate and compare the correctness and response consistency of ChatGPT-5, Gemini 2.5, and Grok 4 on USMLE Step 1-style Allergy/Immunology questions under single and combined prompt conditions.
  • Evaluated 35 USMLE Step 1-style Allergy/Immunology questions across 15 trials per model under two formats: single-question prompts and a combined prompt containing all questions.
  • Measured response variability using Shannon entropy and analyzed the effects of model type, prompt condition, and question difficulty using mixed effects models.
  • Overall accuracy differed significantly (p < 0.001), with Gemini (80.7%) and Grok (80.5%) outperforming ChatGPT (74.3%).
  • Single-item prompts yielded superior accuracy across models, led by Grok (93.1%) and Gemini (90.9%), whereas combined prompts and higher question difficulty significantly reduced performance.
  • Grok demonstrated superior reliability by maintaining the lowest overall response entropy, whereas ChatGPT exhibited the greatest variability across trials.

Cite This Study

Carroll et al. (2026) studied this question.

synapsesocial.com/papers/6a93f0ce6c1a8fb52e79d3dbhttps://doi.org/10.1055/a-2946-7393
View Full Paper
Ask AI
Bookmark
Share

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Comparative Performance of Large Language Models in the Polish State Specialization Examination in Anesthesiology and Intensive Care Medicine2026
  2. 2Performance of 5 AI Models on United States Medical Licensing Examination Step 1 Questions: Comparative Observational Study2026 · 3 citations
  3. 3A Comparative Analysis of Three Large Language Models in Answering Patient Queries on Otolaryngology Emergencies2025
  4. 4Benchmarking five large language models in medical genetics: a bilingual comparative evaluation using published and novel expert-authored questions2026
  5. 5Performance of large language models on the Turkish Pharmacy Specialty Examination: a comparative analysis of accuracy, confidence, and readability2026