PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 14, 2026Scientific Reports0 citationsOpen Access

Evaluating artificial intelligence chatbot performance on board-level geriatrics questions

MZMert ZureMSMetin Sökmen

Key Points

  • This study aims to evaluate the performance of AI language models on board-level geriatrics questions.
  • Tested four AI models on 300 multiple-choice geriatrics questions from BoardVitals.
  • Questions categorized into easy, medium, and hard based on difficulty levels.
  • Each model answered the questions twice to assess accuracy and consistency.
  • Evaluated model responses for alignment with predefined difficulty ratings and explanation quality.
  • GPT-4o had the highest accuracy (85.3%) and consistency (96.3%) among the models tested.
  • All models performed best on easy questions, but accuracy decreased for harder questions (p < 0.001).
  • The agreement between models' difficulty ratings and reference ratings was moderate (mean κ = 0.41).
  • GPT-4o received the highest mean quality score (4.68 ± 0.84), indicating better explanation quality.

Abstract

Artificial intelligence (AI) language models are increasingly being explored as tools to support medical education and clinical care. Evaluating their performance on valid and reliable assessments such as board certification exams may provide insight into their potential integration into real-world medical settings. This study evaluated the accuracy, consistency, and difficulty assessment of four advanced AI models using board-level geriatrics questions. Four AI models-Grok-3, ChatGPT-4o, Microsoft Copilot, and Google Gemini 2.0 Flash-were tested on 300 text-based multiple-choice questions from the BoardVitals geriatrics certification question bank. The questions were equally divided into easy, medium, and hard categories. Each model was asked to classify the question's difficulty and provide an answer twice. Model responses were evaluated for accuracy, consistency between attempts, quality of explanations, and alignment with the difficulty ratings predefined by BoardVitals. GPT-4o demonstrated the highest overall accuracy (85.3%), followed by Grok-3 (82.0%), Copilot (78.7%), and Gemini (74.0%). All models performed best on easy questions, and showed a decrease in accuracy as the difficulty increased (p < 0.001). GPT-4o exhibited the highest consistency (96.3%), followed by Grok-3 (95.0%), Copilot (90.7%), and Gemini (81.3%). While their overall performance surpassed the average success rates of human users in the database, the agreement between model-assigned and reference difficulty ratings was moderate (mean κ = 0.41). GPT-4o received the highest mean quality score (4.68 ± 0.84), followed by Grok-3 (4.59 ± 0.98), Copilot (4.30 ± 1.07), and Gemini (3.88 ± 1.53). Advanced AI models demonstrate strong performance on geriatrics board-level content, suggesting potential applications as educational support tools. However, performance on multiple-choice examinations does not equate to clinical utility. Significant limitations include struggles with complex scenarios, difficulty in metacognitive assessment of question complexity, and variable explanation quality. These findings emphasize that AI integration into geriatric education and practice requires careful human oversight, explicit acknowledgment of limitations, and continued validation in diverse real-world contexts.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zure et al. (2026) studied this question.

synapsesocial.com/papers/69ddd8eee195c95cdefd6780https://doi.org/10.1038/s41598-026-47331-x
Ask AI
Helpful
Bookmark
Share
View Full Paper