Among chatbots, GPT-4 and Bing were the top performers, with Bing performing better at Peru-specific MCQs. Moreover, the educational value of justifications provided by the GPT-4 and Bing could be deemed appropriate. However, it is essential to start addressing the educational value of these chatbots, rather than merely their performance on examinations.
No takes yet. Share an insight, caveat, or question.
Torres-Zegarra et al. (2023) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: