Key points are not available for this paper at this time.
In recent years, the emergence of large language models (LLMs) has transformed the landscape of higher education. Among their most promising applications is the potential to personalise student learning and serve as virtual tutors. This paper investigates whether current versions of the leading LLMs — ChatGPT, Gemini, CoPilot, and Claude — can function as effective learning assistants for accounting students. To this end, we evaluate the accuracy of these models in responding to a set of theoretical and practical multiple-choice questions taken from real university exams. We also examine whether simple prompting strategies, accessible to any student, can improve model performance. Our findings show that while LLMs achieve high accuracy on theoretical questions, their performance declines when solving practical tasks. In particular, accuracy remains insufficient in questions involving the preparation of financial statements. However, when testing more advanced models — ChatGPT-o3 and Claude Opus 4 — we observe a significant leap in performance, suggesting a paradigm shift. These results indicate that, in the near future, accounting students and learners in related fields may be able to reliably use LLMs to support and enhance their independent study.
Llacay et al. (Wed,) studied this question.