Evaluations reveal that registered stability gates may fail under varying response conditions, highlighting implications for data analysis.
Repeated-call psychology studies increasingly use hosted large language models as judges, simulated participants, and classifiers. Some studies preregister a correlation between original and repeated cell rates as an eligibility gate. Correlation, however, targets profile preservation. It can become non-discriminating when reliable between-cell variance is small and can remain high after a common shift. We evaluated a registered Pearson r > .80 gate using a four-model worked case with 14 matched cells, 20 calls per cell, and two runs, together with logistic-normal simulations spanning response baselines, between-cell spreads, and drift mechanisms. The gate excluded the model with the smallest observed mean absolute probability-scale change (.018; λ = .29; r = -.135) and retained the model with the largest (.129; λ = .98; r = .886). Conditional on pooled plug-in cell propensities, a no-drift model in the former response regime failed the registered gate in 83.9% of bootstrap draws. Simulations showed that fixed profile and absolute-drift rules have distinct, baseline-dependent error patterns. The worked case does not establish which model was truly stable. It shows that the registered binary rule could not support a general stability claim across all response regimes. We distinguish mean-propensity, profile, and cell-specific stability and derive requirements for prospective three-outcome calibration. A defensible gate must prespecify its estimand, meaningful drift margin, scale, design, and error budget. Passing should require evidence within the margin, while insufficient information should produce an indeterminate verdict.
No takes yet. Share an insight, caveat, or question.
Boullineau et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: