Comparative evaluation demonstrates deep learning classifies three-tier frailty from wearable gait data in older adults, highlighting scalable objective screening potential.
Background Frailty assessment in older adults typically relies on subjective clinical tools that are time-consuming and require trained personnel, limiting their use in routine or large-scale screening. Wearable sensor-based gait analysis offers a promising objective alternative; however, comprehensive evaluations of deep learning models for multi-class frailty classification using structured gait metrics remain limited. Methods This study evaluates multiple deep learning architectures for three-class frailty classification using structured gait features derived from a single wearable inertial measurement unit (IMU). Four architectures (Transformer, ShapeFormer, InceptionTime, and LSTM-CNN) were evaluated using the publicly available GSTRIDE dataset, comprising 163 older adults (non-frail: n = 65, pre-frail: n = 77, frail: n = 21). Eleven clinically interpretable stride-level gait variables, extracted from a single IMU, were used as model input. A 10-fold participant-level cross-validation scheme, matching the default configuration of the original implementation (cv_seed = 1), was applied to ensure robust performance estimation and prevent data leakage. Results The four architectures yielded comparable mean accuracy under a unified protocol (ShapeFormer 73.1% ± 10.6%, Transformer 72.6% ± 6.0%, InceptionTime 70.5% ± 9.5%, LSTM-CNN 65.9% ± 9.8%). A Friedman omnibus test on the per-fold scores did not reject the null of equal performance (borderline on accuracy and macro-F1 ( p = 0.050 and 0.056 respectively), with no significant difference for macro-AUC [ p = 0.840]), indicating that the apparent ranking is within sampling variability. Performance was higher for non-frail and pre-frail individuals, while classification of frail individuals remained challenging across all models. Conclusion Multi-class frailty classification using structured gait metrics from a single wearable sensor is feasible, though performance varies across models and frailty classes. Model stability and sensitivity to specific classes represent important trade-offs. Significance These findings support the development of scalable, objective approaches to frailty assessment and underscore the importance of selecting models based on clinical priorities.
No takes yet. Share an insight, caveat, or question.
Hughes et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: