Key points are not available for this paper at this time.
Background Artificial intelligence is increasingly integrated into radiology, yet the clinical reliability and generalizability of AI models depend less on algorithmic innovations than on the rigor of preprocessing, training, validation, and performance evaluation. Methodological inconsistencies across these steps remain a major barrier to real-world deployment. Methods This narrative review synthesizes current evidence on methodological practices influencing the development of predictive and segmentation models in radiology. We examine the impact of preprocessing (normalization, resampling, registration, artifacts, annotation), model architecture choice, data splitting, loss function and hyperparameters choice, internal and external validation approaches, and performance metrics. Recent guidelines, including CLAIM 2024, TRIPOD-AI, and Metrics Reloaded, are integrated to frame recommendations. Results The review highlights that preprocessing decisions strongly affect model robustness and cross-center generalization; annotation variability and unrealistic augmentation introduce bias; and inadequate separation of training, validation, and test data leads to misleading performance estimates. Internal validation methods such as cross-validation and bootstrap improve reliability, whereas external and temporal validation remain essential for clinical translation. We show that no single performance metric is sufficient: confusion matrix–based measures, AUROC/AUPRC, overlap scores, boundary metrics, topological descriptors, and lesion-level metrics each capture different facets of performance and must be applied in a task-specific, complementary manner. Conclusions Effective optimization of radiologic AI models requires a comprehensive methodological framework spanning preprocessing, model development, validation, and metric selection. Transparent reporting and adherence to established guidelines are fundamental to ensuring reproducibility, generalizability, and clinical applicability. A structured, method-oriented approach is essential for transforming AI tools into trustworthy solutions in radiology.
Chierici et al. (Fri,) studied this question.