PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 25, 2026BMC Medical Education2 citationsOpen Access

Consistency over accuracy: run-to-run stability of contemporary large language models on Turkish curriculum-aligned theoretical anatomy multiple-choice questions

View Full Paper
ÖGÖmer Alperen GürsesİCİsmail Ceylan

Key Points

  • This research aims to evaluate the accuracy and stability of large language models on anatomy multiple-choice questions in Turkish.
  • Analyzed responses of contemporary large language models on curriculum-aligned anatomy questions.
  • Evaluated single-trial accuracy versus run-to-run stability metrics.
  • Assessed consistent-correct and consistent-wrong rates for stability appraisal.
  • Models demonstrated high accuracy on individual trials but exhibited substantial volatility.
  • Run-to-run stability varied, indicating systematic errors not visible in single trials.
  • Recommendations include prioritizing stability in the selection of models for educational use.

Abstract

In Turkish, curriculum-aligned anatomy items, contemporary LLMs can be both accurate and stable, but single-trial accuracy can mask volatility and stable systematic errors. Adoption decisions should prioritize stability-aware appraisal (including consistent-correct and consistent-wrong rates), with local validation on institutional item banks and periodic re-evaluation as models evolve. Extending this framework to multimodal anatomy and constructed-response tasks will further inform trustworthy, learner-facing use.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Gürses et al. (2026) studied this question.

synapsesocial.com/papers/6975b350feba4585c2d6ebd9https://doi.org/10.1186/s12909-026-08656-3
Ask AI
Helpful
Bookmark
Share
View Full Paper