Key points are not available for this paper at this time.
Accurate skin lesion classification is hard because classes can look similar, datasets are imbalanced, and devices and domains vary. We introduce SynthraXCoreNet, a six CNN ensemble with ResNet50V2, ResNet101V2, ResNet152V2, DenseNet201, NASNetLarge, and Xception. Each backbone is trained per dataset, and at inference we average posterior probabilities from a shared 224×224×3 input. We assess probability calibration with reliability diagrams and Expected Calibration Error using temperature scaling. For interpretability we combine Integrated Gradients, Grad CAM variants, and saliency maps, and we score explanation quality by faithfulness, robustness, complexity, and axiomatic metrics. We include a small ablation that contrasts equal soft voting with weighted soft voting on Accuracy, Macro F1, and ECE. This work targets clinical decision support rather than autonomous diagnosis. The current scope is single lesion dermoscopic images. Multi center dermatologist reviewed prospective validation and extensions to multi lesion use are in progress. On four public datasets, test performance is reported as mean ± 95% Wilson CI. On HAM10000, Accuracy is 98.85% ± 0.68 and Macro F1 is 0.99. On ISIC2019, Accuracy is 95.30% ± 0.83 and Macro F1 is 0.95. On BCN20000, Accuracy is 95.00% ± 1.21 and Macro F1 is 0.95. On DERM12345, Accuracy is 96.63% ± 1.00 and Macro F1 is 0.97. Confusion analyses and calibration checks are consistent with strong accuracy and well calibrated probabilities. End to end inference latency is ≈ 461 ms per image with batch 1 on an RTX 4070 Ti in FP32, which supports interactive decision support at about 2.17 images per second or about 130 images per minute.
Khan et al. (Mon,) studied this question.