Concept Bottleneck Models (CBMs) aim to improve interpretability by forcing predictions to pass through human-interpretable concepts. However, many CBMs achieve high predictive accuracy by bypassing the bottleneck through latent shortcut features, a phenomenon known as concept leakage. Such behavior weakens the reliability of concept-based explanations, particularly in high-stakes clinical applications. To address this, we propose LCBM (Leakage-Constrained Bottleneck Model), a purified concept bottleneck architecture that decomposes latent representations into two structurally decorrelated subspaces: a concept-aligned semantic space (zₒ₄₌) containing clinically meaningful features and a residual space (zₑ₄ₒ) capturing non-diagnostic artifacts. Through orthogonality regularization and an adversarial Gradient Reversal Layer (GRL), the framework minimizes the propagation of diagnostic information through non-semantic pathways. The framework is evaluated on the Derm7pt and HAM10000 dermatology benchmarks using ResNet and Vision Transformer backbones. Our results demonstrate that LCBM significantly improves the semantic reliability of CBMs, reducing the leakage gap from 0. 4487 to 0. 0539 under ground-truth concept intervention. The model achieves competitive AUC scores of 0. 8358 (ResNet-50) and 0. 8107 (ViT-B/16) on Derm7pt, while reaching an AUC of 0. 94 on HAM10000. These results suggest that structured latent purification through decorrelation can enhance faithful concept-based reasoning in clinical AI systems, providing a more transparent foundation for diagnostic decision support.
Kandpal et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: