This research aims to evaluate the performance of large language models (LLMs) in emergency medicine specialty examinations using the TUS format.
Cross-sectional study design comparing LLM performance to traditional examination formats.
Analyzed responses to TUS questions in the context of linguistic and curricular differences.
Focused on emergency medicine specialty examinations to assess contextual relevance.
LLMs demonstrated varying levels of performance when answering TUS questions compared to USMLE-style exams.
Performance metrics indicated strengths and weaknesses specific to the emergency medicine context, influencing assessment validity.
Abstract
Unlike prior studies primarily focused on USMLE-style examinations, this study evaluates LLM performance using the TUS, which reflects a different linguistic and curricular context.By directly comparing LLMs