Abstract Purpose To quantitatively compare end-to-end training of a convolutional neural network (CNN) with transfer learning using a frozen VGG16 feature extractor for multiclass brain tumor classification on contrast-enhanced magnetic resonance (MR) images. Methods A publicly available dataset of 6, 056 contrast-enhanced T1-weighted brain MRI images (glioma, n = 2, 004; meningioma, n = 2, 004; and brain tumor, n = 2, 048) was preprocessed (resized to 224 × 224, intensity-normalized, pseudo-RGB) and stratified (70/15/15 split). Five independent runs with distinct random seeds (42, 123, 2024, 7, 999) compared (1) a custom CNN trained end-to-end and (2) a frozen VGG16 with a task-specific head under identical conditions. Performance was assessed via accuracy, precision, recall, F1 score, ROC-AUC, and generalization gap (Δgen = training–testing accuracy) with paired t tests and Cohen’s d. Results The frozen VGG16 model achieved a mean test accuracy of 94. 4% ± 0. 4% (95% CI ± 0. 35%), significantly outperforming the baseline CNN (58. 8% ± 25. 0%; paired t test p = 0. 032; Cohen’s d = 1. 45). VGG16 showed a near-zero generalization gap (Δgen = − 0. 022) versus severe overfitting at baseline (Δgen = + 0. 344). Conclusion Under strictly controlled single-dataset conditions, frozen VGG16 feature extraction establishes a statistically robust and reproducible baseline that demonstrates significant superiority over end-to-end training of a lightweight CNN (mean test accuracy 94. 4% ± 0. 4% vs. 58. 8% ± 25. 0%; paired t test p = 0. 032; Cohen’s d = 1. 45). External multicenter validation on diverse scanners and populations, together with quantitative interpretability analyses, remain essential prerequisites for clinical translation.
Al-Sarairah et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: