Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 30, 2026Applied Clinical Informatics

A Comparative Analysis of Large Language Model Performance on USMLE Step 1-Style Allergy/Immunology Questions: Evaluating Correctness and Consistency

View Full Paper
Ask AI
Bookmark
Share

Authors

MCM. CarrollSKSabrina KentisHKHannah Kareff

Discussion

Loading...

Member takes

Overview

Benchmarking study demonstrates variable accuracy and consistency across prompt formats on medical licensing questions, highlighting prompt sensitivity in AI education.

Key Points

  • To evaluate and compare the correctness and response consistency of ChatGPT-5, Gemini 2.5, and Grok 4 on USMLE Step 1-style Allergy/Immunology questions under single and combined prompt conditions.
  • Evaluated 35 USMLE Step 1-style Allergy/Immunology questions across 15 trials per model under two formats: single-question prompts and a combined prompt containing all questions.
  • Measured response variability using Shannon entropy and analyzed the effects of model type, prompt condition, and question difficulty using mixed effects models.
  • Overall accuracy differed significantly (p < 0.001), with Gemini (80.7%) and Grok (80.5%) outperforming ChatGPT (74.3%).
  • Single-item prompts yielded superior accuracy across models, led by Grok (93.1%) and Gemini (90.9%), whereas combined prompts and higher question difficulty significantly reduced performance.
  • Grok demonstrated superior reliability by maintaining the lowest overall response entropy, whereas ChatGPT exhibited the greatest variability across trials.

Cite This Study

Carroll et al. (2026) studied this question.

synapsesocial.com/papers/6a93f0ce6c1a8fb52e79d3dbhttps://doi.org/10.1055/a-2946-7393
View Full Paper
Ask AI
Bookmark
Share