Key points are not available for this paper at this time.
Effect sizes are widely used in psychological science, but with a focus solely on mean differences. This study compares three mean-based indices: Cohen's d , the Common Language Effect Size (CLES), the parametric overlap coefficient derived from d (η p ), to the non-parametric Overlapping Index (η), which focuses on the entire distributional differences. Using skew-normal distributions to generate controlled scenarios of violations of normality and variance homogeneity, and symmetry, we systematically manipulated mean differences, variance ratios, skewness, and sample size to evaluate each index in terms of Relative Mean Bias, Normalized Root Mean Square Error, and 95% Coverage. Cohen's d , CLES, and η p showed near-perfect descriptive correlations; however, their performance differed substantially. Cohen's d consistently exhibited low bias, high precision, and accurate coverage across all scenarios, whereas CLES and η p showed substantial bias and low coverage, particularly under skewness and heteroscedasticity. The non-parametric index η remained unbiased under shape differences and variance heterogeneity but performed less reliably when the populations truly overlapped. The results indicated that effect-size indices derived from d are not interchangeable, and that high empirical correlations do not guarantee same precision. Cohen's d remains the most robust estimator of a location difference, whereas, η provides a more complete view of the entire distribution. Here, we argue that researchers should select effect sizes based on their statistical properties and propose a shift toward interpreting effects in light of the full distribution, rather than through mean-based conventions.
Perugini et al. (Tue,) studied this question.