Researchers increasingly must judge whether a hosted large language model (LLM) has held its behaviour steady across a study window. A natural response is to re-run selected conditions and apply a test-retest correlation threshold. We show that this approach can become uninformative or inverted when responses sit near a ceiling or floor. With responses tightly clustered, matched conditions hold too little variance, so the correlation only tracks sampling noise, and a perfectly stable model can be excluded. The threshold has a mirror-image flaw. When a model's propensity shifts uniformly across conditions, rank order holds and a drifting model can pass. Across an empirical study of frontier models and systematic simulations, the correlation rule can turn a measurement limitation into an incorrect exclusion. The problem is not specific to LLMs. Any screen that gates on a reliability coefficient near a response ceiling or floor can fail the same way. We offer a three-outcome screen that asks whether the overall run shifted, whether the design holds enough reliable between-condition information, and whether particular conditions moved. It returns a pass, fail, or indeterminate stability verdict, requiring calibration before data collection and declining to rule where the design offers too little variance to earn one.
Boullineau et al. (Fri,) studied this question.