Smartwatches with movement sensors have emerged as promising tools for monitoring dietary behavior. This study uses commercial smartwatch sensor data, namely tri-axial acceleration together with device-derived orientation (pitch and roll) and power features, to detect eating episodes, with rigorous subject-independent validation. Twenty healthy participants were equipped with smartwatches and video recorded during four meals under semi-naturalistic settings. Raw sensor data was transformed using time-windowed features. Multiple machine learning and deep learning methods were evaluated using Leave-One-Subject-Out (LOSO) cross-validation. Over 26,000 Information Units from 19 subjects were analyzed. Using only motion-derived predictors, XGBoost with hyperparameter tuning achieved the best balanced accuracy (0.639 95% CI 0.607, 0.668) and AUC (0.699 0.653, 0.743); the full feature set that additionally used experimental-context variables (meal, food and menu information) gave near-identical performance, and paired cluster-bootstrap comparisons indicated that Logistic Regression remained close on balanced accuracy. A Transformer encoder achieved the numerically highest sensitivity (0.691 0.637, 0.735) at a significantly lower specificity (0.504 0.456, 0.549); under nested threshold selection and multiplicity-adjusted comparison the sensitivity advantage over tuned XGBoost was not statistically significant. We demonstrate that eating detection from smartwatch sensors remains challenging when evaluated with proper subject-independent validation. The gap between within-subject and between-subject performance reflects high inter-individual variability in eating gestures, a key limitation for real-world deployment.
Vedovelli et al. (Tue,) studied this question.