Key points are not available for this paper at this time.
Abstract It has been documented that spread–error equality and a flat rank histogram are necessary but insufficient for demonstrating ensemble forecast reliability. Nevertheless, these metrics are heavily relied upon—both in the literature and at operational numerical weather prediction centers—as if they were valid indicators of perfect ensemble dispersion. In this study, we show that, in the presence of a climatological variance bias, several reliability diagnostics can indicate falsely that forecasts are perfectly reliable. This is demonstrated theoretically for the spread–error relationship and related reliability budget. Idealized experiments based on the multivariate normal distribution reveal that the same joint‐distribution structure causing insufficiency also leads to false diagnoses of reliability with the rank histogram and the reliability component of the continuous rank probability score. Under this structure, and when the ensemble‐mean state is meaningfully different from climatology, the truth lies systematically among the least extreme members when climatological variance is excessive in each member, and among the most extreme members when climatological variance is deficient. Importantly, this behavior is also shown to be plausible in operational ensemble weather forecasts. Combining these results with calibration principles from statistical postprocessing leads us to conclude that both “perfect dispersion” and “underdispersion” are ill‐defined. When diagnostics are misinterpreted as indicating the latter, improper tuning of forecasts can lead to further deterioration of forecast quality, even while improving spread–error and rank histogram behavior. To address these issues, we propose a reliability diagnostic based on three easily computed statistics, directly related to the structure of the joint distribution of ensemble members and the reference truth up to second order. The diagnostic separates contributions to unreliability originating from climatology and predictability, enabling a more precise and robust characterization of ensemble behavior.
Dirkson et al. (Fri,) studied this question.