PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 16, 2026Diseases0 citationsOpen Access

Multi-Task Deep Learning Model for Automated Detection and Severity Grading of Lumbar Spinal Stenosis on MRI: Multi-Center External Validation

View Full Paper
PUPhatcharapon UdomluckWCWatcharaporn CholamjiakJIJakkaphong Inpun

Key Points

  • Evaluate the effectiveness of deep learning methods for grading lumbar spinal stenosis in MRI scans.
  • Processed axial MRI images using pretrained VGG19, ConvNeXt-Tiny, and DINOv2 models
  • Trained logistic regression, support vector machine, and LightGBM classifiers
  • Conducted external validation on MRI data from the University of Phayao Hospital
  • Assessed performance with accuracy, precision, recall, F1-score, and ROC curves
  • VGG19-based features achieved the highest accuracy of 0.9556 and F1-score of 0.9558
  • External validation showed AUC values between 0.994 and 1.000 across severity grades
  • SVM and LightGBM also performed well with accuracies of 0.9333 and 0.9222, respectively
  • DINOv2 exhibited reduced generalizability, particularly with LightGBM (accuracy 0.6222)

Abstract

Background/Objectives: Accurate and reproducible grading of lumbar spinal stenosis (LSS) is clinically critical for guiding treatment decisions and patient management, yet manual assessment remains challenging due to imaging variability and inter-observer subjectivity. To address these limitations, this study aimed to evaluate the generalizability of deep learning–based feature extraction methods—VGG19, ConvNeXt-Tiny, and DINOv2—combined with classical machine learning classifiers for automated multi-grade LSS assessment. Automated grading enables objective, reproducible, and scalable assessment of lumbar spinal stenosis severity, addressing key limitations of manual interpretation. Methods: Axial MRI images were processed using pretrained VGG19, ConvNeXt-Tiny, and DINOv2 models to extract deep features. Logistic Regression, Support Vector Machine (SVM), and LightGBM were trained on internal datasets and externally validated using MRI data from the University of Phayao Hospital. Performance was assessed using accuracy, precision, recall, F1-score, confusion matrices, and multi-class ROC curves. Results: VGG19-based features yielded the strongest external performance, with Logistic Regression achieving the highest accuracy (0.9556) and F1-score (0.9558). External validation further demonstrated excellent discrimination, with AUC values ranging from 0.994 to 1.000 across all severity grades. SVM (0.9333 accuracy) and LightGBM (0.9222 accuracy) also performed well. ConvNeXt-Tiny showed stable cross-model performance, while DINOv2 features exhibited reduced generalizability, especially with LightGBM (accuracy 0.6222). Most classification errors occurred between adjacent grades. Conclusions: Deep convolutional features—particularly VGG19—combined with classical machine learning classifiers provide robust and generalizable LSS grading across external MRI data. Despite advances in modern architectures, CNN-based feature extraction remains highly effective for spinal imaging and represents a practical pathway for clinical decision support.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Udomluck et al. (2026) studied this question.

synapsesocial.com/papers/6969d4fd940543b977709efbhttps://doi.org/10.3390/diseases14010032
Ask AI
Helpful
Bookmark
Share
View Full Paper