Empirical evaluation demonstrates class-conditional calibration restores minority coverage under severe imbalance in tabular models, indicating risks when calibration samples are scarce.
Split conformal prediction provides distribution-free marginal coverage guarantees, but marginal coverage can mask severe under-coverage of minority classes when the underlying label distribution is imbalanced. Class-conditional (Mondrian) conformal calibration is the standard remedy, and recent largescale benchmarks report that it reliably restores minority-class coverage in high-stakes, severely imbalanced settings. What remains comparatively under-examined is whether this correction is uniformly beneficial, or whether its value depends jointly on the severity of the imbalance and on the number of calibration examples available per class. This study investigates that question empirically across four widely used tabular model families — logistic regression, random forest, histogram-based gradient boosting, and a radial-basis-function support vector machine — on two public datasets under three imbalance regimes: a mildly imbalanced binary dataset, a severely imbalanced binary dataset obtained by controlled minority-class subsampling, and a mildly imbalanced three-class dataset with small per-class calibration samples. Over 30 independent random data splits per condition, class-conditional calibration reduced the maximum-minimum per-class coverage gap relative to marginal calibration by a large, statistically significant margin under severe binary imbalance (mean gap reduction of 0.16–0.18 across models, Holm-adjusted p < 0.01 for all four model families), consistent with prior reports. Under mild binary imbalance the correction produced no significant fairness benefit while still inflating the average prediction-set size. Under mild three-class imbalance with small per-class calibration samples, class-conditional calibration significantly increased the coverage gap in three of four model families (Holm-adjusted p < 0.05), reversing its intended effect. These results indicate that class-conditional conformal calibration is not a uniformly safe default: its benefit is concentrated in regimes of severe imbalance with adequate per-class calibration data, and it can be counterproductive when classes are numerous, close to balanced, or calibration data are scarce. The study contributes a joint, cross-model characterization of this trade-off and provides a fully reproducible experimental pipeline for auditing the fairness-efficiency behavior of conformal calibration schemes before deployment.
No takes yet. Share an insight, caveat, or question.
Ahmad Raza (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: