PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 2, 2026BMC Research Notes2 citationsOpen Access

Evaluating the ability of AI models to generate level-specific medical MCQs with variable difficulty

View Full Paper
MAManar Al-lawamaOAOmar AltamimiEAEyad Altamimi

Key Points

  • Evaluate the effectiveness of AI in generating medically relevant multiple-choice questions across different difficulty levels.
  • Conducted a mixed-methods study
  • Used structured prompts for ChatGPT to generate MCQs for clinical-year medical students
  • Employed iterative refinements with expert assessments of clarity and accuracy.
  • 87% of items rated high for content accuracy
  • 82% of items showed clarity
  • Only 30% of distractors were plausible

Abstract

Artificial intelligence (AI), particularly large language models (LLMs) such as ChatGPT, is increasingly applied in medical education to automate assessment design. However, concerns persist regarding the content accuracy, cognitive depth, and psychometric validity of AI-generated multiple-choice questions (MCQs). A mixed-methods study was conducted to evaluate a structured prompt guiding ChatGPT in generating clinically relevant, single-best-answer MCQs in pediatrics. The prompt defined item count, subdomain distribution, difficulty, and examinee level. Following seven iterative refinements, ChatGPT produced 100 MCQs targeting clinical-year medical students. Two blinded experts independently assessed each item for clarity, content accuracy, clinical realism, distractor plausibility, and cognitive alignment. All items were structurally coherent and linguistically sound. Content accuracy was rated high in 87% of items, and stem clarity in 82%. Distractor plausibility was acceptable in 77%, though 23% of items contained at least one implausible distractor. Agreement between AI-predicted and expert-rated difficulty was low (κ = 0.06), suggesting limited calibration.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Al-lawama et al. (2026) studied this question.

synapsesocial.com/papers/69800910aa6434d8c2036e3bhttps://doi.org/10.1186/s13104-026-07683-z
Ask AI
Helpful
Bookmark
Share
View Full Paper