Summary As large language models (LLMs) become increasingly integrated into analytical workflows, an urgent question arises: Can these models replace the trained statistician? This paper presents a controlled experiment to directly test the statistical reasoning ability of LLMs. We employ Monte Carlo simulations to generate datasets with known ground‐truth parameters and pose four canonical statistical questions to five commercially prominent models across three linguistically distinct prompt formulations and five sampling temperature settings, yielding 3000 observations in a full‐factorial design. In order to find an answer to our question, we design prompts that reflect how decision‐makers with varying degrees of statistical knowledge would query AI in the absence of a trained statistician. We find that prompts that portray higher statistical competency can result in higher accuracy for some (but not all) LLMs; we also find that LLMs can fail catastrophically on tasks requiring quantitative precision. We connect these findings to architectural differences among models and to recent literature on epistemic mirroring in LLMs and argue that the observed patterns reveal models are performing linguistic pattern matching on statistically flavoured text rather than genuine statistical reasoning. We conclude that current LLMs cannot replace the statistician, though certain architectures approach useful performance on pattern‐recognition subtasks.
Jank et al. (Sun,) studied this question.