Benchmarking study demonstrates variable accuracy and consistency across prompt formats on medical licensing questions, highlighting prompt sensitivity in AI education.
Key Points
To evaluate and compare the correctness and response consistency of ChatGPT-5, Gemini 2.5, and Grok 4 on USMLE Step 1-style Allergy/Immunology questions under single and combined prompt conditions.
Evaluated 35 USMLE Step 1-style Allergy/Immunology questions across 15 trials per model under two formats: single-question prompts and a combined prompt containing all questions.
Measured response variability using Shannon entropy and analyzed the effects of model type, prompt condition, and question difficulty using mixed effects models.
Overall accuracy differed significantly (p < 0.001), with Gemini (80.7%) and Grok (80.5%) outperforming ChatGPT (74.3%).
Single-item prompts yielded superior accuracy across models, led by Grok (93.1%) and Gemini (90.9%), whereas combined prompts and higher question difficulty significantly reduced performance.
Grok demonstrated superior reliability by maintaining the lowest overall response entropy, whereas ChatGPT exhibited the greatest variability across trials.