Large language model-based chatbots demonstrated variable performance in answering ophthalmology questions at the medical school level. Among the models evaluated, Microsoft Copilot performed best in both accuracy and readability. These findings suggest that model choice may influence the usefulness of AI-generated content in educational settings.
Çakmak et al. (Wed,) studied this question.