Key points are not available for this paper at this time.
Artificial intelligence (AI) has shown strong potential in dermoscopic image classification; however, reliable clinical deployment remains challenging due to limited generalizability across populations and institutions. This challenge is further compounded by the performance saturation of single models: continued architectural scaling often results in diminishing returns within the training domain and fails to yield corresponding gains in cross-domain robustness. Previous work often relies on a single benchmark dataset and evaluate models under matched data distributions, which can overestimate real-world performance. In this study, we examine the robustness and generalizability of dermoscopic AI systems under realistic deployment conditions. We curate HAM20000, a controlled expansion of the HAM10000 dataset integrating multiple public ISIC releases from 2017 to 2024, designed as a stress-test dataset to analyze performance saturation and architectural sensitivity. A total of 27 pretrained deep learning models, including convolutional neural networks, Vision Transformers, and hybrid CNN–Transformer architectures, were systematically evaluated. To mitigate the limitations of single-model optimization, we applied greedy ensemble selection to construct a compact heterogeneous ensemble using soft voting. All models were trained and selected exclusively on HAM20000 and evaluated under strict zero-shot external validation on two independent datasets: CSMUH (East Asian population) and BCN20000 (European cohort). Although the proposed ensemble achieved over 93% accuracy on the HAM20000 dataset, its performance declined substantially under external validation, highlighting a pronounced cross-population domain gap. On the BCN20000 external validation set, the ensemble reached 62% accuracy, outperforming the strongest individual model (59%) by 3%. On the CSMUH external validation set, the proposed GES achieved an accuracy of approximately 57%, demonstrating improved performance over representative single-model baselines, although it did not exceed the best-performing individual architecture in this dataset. These results indicate that, despite a marked absolute performance drop under distribution shift, the ensemble provides a modest yet consistent improvement over single-model baselines rather than a substantial performance leap. This finding underscores the importance of rigorous external validation and suggests that heterogeneous ensemble learning represents a practical and promising strategy for achieving incremental robustness gains in clinically applicable dermoscopic decision support systems.
Tseng et al. (Sun,) studied this question.