Abstract. Repeated calls to a hosted large language model (LLM) are often analysed as though each output were an independent participant. That shortcut usually gets the unit of evidence wrong. Calls sent under the same prompt, model version, decoding settings, safety layer, and collection window are better understood as repeated observations from a prompt-by-condition cell. This paper introduces a preregisterable five-module audit protocol for repeated-call behavioural LLM studies. It covers cell definition, model-specific calibration, response taxonomy, endpoint stability, and reviewer-verifiable provenance. A Monte Carlo demonstration shows why the distinction matters. Naive call-level analyses produced false-positive rates approximately two to six times those of the dispersion-corrected comparator when cells differed in their underlying response tendencies. Aggregating calls to cell counts did not fix the problem when the variance model still treated dispersion as fixed. Dispersion-corrected cell-level analysis and a small-sample corrected covariance analysis were near nominal in the simulated setting. We then use the protocol in a preregistered four-model moral-judgement experiment with 11,200 trials and 40,871 logged records. The audit identified ceiling-level stimuli, model-specific parser failures, contaminated controls, and a stability gate that excluded one near-ceiling model despite limited absolute drift. Reusable templates, simulation code, figures, and hash-locked artefacts support reviewer verification. The protocol gives prompt engineering, alignment evaluation, bias auditing, cognitive testing, and social-judgement research a route for turning repeated hosted outputs into bounded evidence. Deposit contents. This Zenodo version contains the preprint PDF only. Related OSF registration: https://osf.io/6m3sk/. Version note. Preprint v3.1 adds an introduction signpost that positions the protocol relative to concurrent statistical work on the non-independence of repeated model generations, and adds page numbers. It builds on v3, which added the underlying citations (a NIST technical report on statistical models for AI evaluation, Keller et al., 2026; and peer-reviewed benchmark work, Zhang et al., 2026, Findings of the ACL; Song et al., 2025, NAACL), redrew the protocol figure, and improved table presentation. No analyses, numerical results, data, or conclusions were changed.
Boullineau et al. (Mon,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: