Key points are not available for this paper at this time.
The latest advancements in artificial intelligence (AI) tools have sparked controversy among educators about their usage in statistics classes. This study examines the capabilities and limitations of two popular large language models (LLMs), namely ChatGPT and Gemini, for learning statistics. A comparative approach was used to examine the accuracy and consistency of responses from LLMs regarding selected descriptive statistical concepts, and whether they can support concept-based learning or merely respond to queries. Twenty-seven purposively selected problems covering central tendency and dispersion were presented to both LLMs. Problems included 20 routine tasks (8 direct calculations and 12 multiple-choice questions), 5 non-routine problems requiring multi-step reasoning, and 2 special cases that tested error detection and information gap identification. The responses were assessed using an evaluation criterion that measures computational accuracy, procedural consistency across verification prompts, and quality of conceptual interpretation; two assessors independently rated the responses and resolved the disagreements. Results showed that both LLMs demonstrated similar high accuracy. However, ChatGPT demonstrated superior performance, with notable advantages in consistency and conceptual learning support. While both LLMs exhibited competence on all computational tasks, ChatGPT provided more detailed step-by-step explanations proactively, maintained stable responses when verification prompts were issued, and successfully detected misinformation and missing information in all cases. Gemini showed strengths in initial method selection but exhibited critical weaknesses, including procedural inconsistency, interpretation errors, and complete failure to detect problematic information even after multiple prompts with hints. Additionally, ChatGPT produced half the errors of Gemini, with Gemini's errors more frequently involving conceptual misunderstandings. Findings indicate that LLM educational value depends critically on pedagogical coherence beyond computational accuracy. Educators should be guided by best practices for LLMs' integration in statistics education, ensuring these tools support rather than substitute for the learning process.
Almarashdi et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: