Key points are not available for this paper at this time.
Radiomics has produced tens of thousands of publications yet almost no handcrafted radiomic signatures in routine clinical use, and the reasons are increasingly understood to be problems of reproducibility and clinical translation rather than of algorithms. This critical narrative review argues that the field systematically generates paper-grade evidence—findings sufficient to publish—far faster than decision-grade evidence—findings sufficient to change clinical practice. Drawing on meta-scientific research, we describe seven fragility mechanisms (publication bias, analytical flexibility, underpowering, HARKing hypothesizing after the results are known, citation distortion, cognitive bias, and misaligned incentives) and show why radiomics is structurally exposed to all of them simultaneously: high-dimensional feature spaces, acquisition-dependent measurement instability, segmentation variability, retrospective single-centre data, small samples, and leakage-prone validation. We then summarise empirical evidence on the radiomics literature, which remains pervaded by suboptimal methodological quality, near-absent negative results, limited external validation, sparse calibration and clinical-utility assessment, low data and code sharing, and a measurable retraction signal. We interpret these patterns as the output of a self-reinforcing system rather than isolated errors, and argue that better algorithms alone cannot resolve them. Finally, we argue that closing this gap requires not better models but evidentiary discipline: the consistent, enforceable application of standards the field already has, and the calibration of published claims to the strength of the underlying evidence.
Pozzi et al. (Mon,) studied this question.