Randomized trial evaluates sleep tracker accuracy in home settings, indicating reliable performance in monitoring.
Introduction Consumer sleep trackers (CSTs) are increasingly used for daily sleep monitoring, yet their performance in real-world home environments remains insufficiently characterized. Most existing validations have been conducted in controlled laboratory settings, which may not fully reflect everyday sleeping conditions. This study aimed to evaluate the accuracy of five common CSTs, including two smartwatch-based, two ring-based, and one airable device, by directly comparing their measurements against simultaneously collected home polysomnography (PSG) in participants’ habitual sleep environments. Methods Eight adults participated, each undergoing overnight monitoring in their typical home setting while minimizing external noise and ensuring single-sleeper conditions. Five CSTs (Apple Watch, Galaxy Watch, Galaxy Ring, Oura Ring, and Sleep Routine) were worn concurrently along with a type 2 home PSG device (Embletta MPR PG & ST+Proxy). Sleep stages were evaluated using epoch-by-epoch comparison across 9,441 epochs (78.7 hours) with four-stage classification (Wake, Light, Deep, REM). Sleep onset and wake time accuracy were also assessed using mean bias and mean absolute error (MAE). Results For 4-stage sleep classification, Sleep Routine demonstrated the highest macro F1 score (0.83), followed by Apple Watch (0.79), while the lowest-performing device scored 0.71. Accuracy showed a similar pattern, with Sleep Routine at 86.5%, Apple Watch at 84.0%, and the lowest device at 75.1%. For sleep onset time, Galaxy Watch exhibited the smallest mean bias (-0.2 min) but with larger variability, whereas Sleep Routine and Apple Watch achieved the lowest MAE (2.2 min). For wake time estimation, Sleep Routine achieved near-zero mean bias (< 0.1 min) and the lowest MAE (1.0 min). Conclusion In real-world home environments, all CSTs demonstrated generally high performance, with every device achieving a macro F1 score above 0.70. These findings indicate that CSTs can achieve reliable performance under naturalistic yet controlled home conditions, while also showing performance differences across device types and sensing modalities. Larger and more diverse studies will further enhance the generalizability of these results. Support (if any)
No takes yet. Share an insight, caveat, or question.
Kim et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: