Shapelet-based classifiers offer structural interpretability: discriminative subsequences form an inspectable vibration-pattern vocabulary and a shallow decision-tree ensemble produces traceable fault-type predictions. We impose explicit interpretability constraints on the model and apply Bayesian optimisation within this bounded region; the primary experimental question is whether these constraints carry a performance penalty relative to an unconstrained baseline. Evaluation uses recording-level cross-validation on CWRU (Case Western Reserve University) and MFPT (Machinery Failure Prevention Technology) bearing datasets with Hilbert envelope demodulation. The central finding is that the constraints impose no systematic performance penalty: the shapelet classifier matches ROCKET, a non-interpretable baseline, on both datasets, with cross-validated mean F1 differences smaller than the fold-to-fold standard deviation. To further characterise the selected models under the controlled laboratory conditions studied here, we assess probability calibration and conformal prediction coverage as secondary analyses. Raw probability estimates are well-calibrated, but Platt scaling degrades under cross-severity distribution shift; split conformal prediction yields valid coverage on CWRU but fails on MFPT due to class-proportion mismatch across recording-level splits. Together, these results show that structural constraints supporting interpretability are compatible with competitive performance, and identify the conditions under which reliability tools succeed and fail in this setting.
Gonzalez-Garcia et al. (Fri,) studied this question.