Key points are not available for this paper at this time.
Consumer-grade wearable sensors may enable continuous monitoring of pilot workload and stress during flight training, yet most prior studies rely on simulators, raw-score labelling, and within-subject validation, limiting generalisability. This study evaluates whether electrodermal activity (EDA), electrocardiogram (ECG)-derived features, and wrist skin temperature, recorded from an Empatica Embrace Plus and a Polar H10 during real Cessna 172 flight training, can classify pilots’ task-relative workload and stress deviations. Thirty-five pilots completed four flight segments and rated workload and stress after each. Fold-safe two-way residual binary labels removed inter-pilot scale-use differences and task-level effects, and five classifiers were evaluated under leave-one-subject-out (LOSO) cross-validation with Benjamini–Hochberg FDR correction. Under LOSO, a Linear SVC on combined features classified stress (macro F1 = 0.607) and XGBoost on EDA classified workload (macro F1 = 0.598) significantly above chance (padj=0.033); both remained stable under nested cross-validation with an inner hyperparameter search (nested 0.606 and 0.561). A LightGBM model on EDA gave a numerically higher stress score (0.611) that did not survive nested validation. Subject-dependent within-subject validation produced higher apparent performance (macro F1 = 0.853 for stress and 0.791 for workload), but a stricter within-pilot analysis was unstable. These contrasts indicate that personalised classification may be feasible after calibration, whereas uncalibrated cross-pilot prediction in real flight remains modest, with post-flight debriefing the most plausible near-term application.
Xu et al. (Sat,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: