Abstract Introduction Sound-based artificial intelligence (AI) models for sleep staging offer a non-contact approach to sleep monitoring. Prior studies have relied mainly on data collected from a Korean clinical sleep center, where the patient population was predominantly Asian. While effective regionally, the geographic specificity of training data raises questions about global applicability. External validation across diverse clinical environments and scoring standards is therefore essential, and this study addresses that need by evaluating the generalizability of a sound-based sleep staging model using a U.S. clinical cohort. Methods Fourteen synchronized PSG audio recordings were collected at Stanford Sleep Medicine Center. During PSG acquisition, sleep sounds were simultaneously recorded on an iPhone at the bedside, and the model inferred sleep stages solely from smartphone audio. The final cohort included 14 participants (9 females, 5 males; mean age 57.5 ± 18.5 years; BMI 29.6 ± 6.1; range 25–78). Racial and ethnic distribution was White-Hispanic (n=3), White-Non-Hispanic (n=9), Other-Hispanic (n=1), and Asian (n=1). Epoch-by-epoch accuracy and macro F1 were calculated to assess performance on this external dataset. Results The model was trained initially on 2,973 nights of synchronized PSG–audio data from a Korean clinical sleep center (SNUBH) and achieved an accuracy of 80.9% and macro F1 0.77 on an independent test set of 802 subjects (mean age 50.5 ± 15.4; BMI 25.5 ± 3.6; male:female = 530:272; all Asian). External validation using U.S. smartphone-recorded audio demonstrated an accuracy of 77.5% and macro F1 of 0.74, indicating consistent performance despite geographic and scoring differences. Notably, the overall performance degradation was marginal, with a drop of 0.04 in F1 score across the four sleep stages. Conclusion This pilot external validation confirms the generalizability of our sound-based AI model, demonstrating its potential adaptability to diverse racial and anthropometric profiles beyond its original Korea-based dataset of primarily Asian participants. The findings highlight the feasibility of smartphone-based, non-contact sleep staging and strengthen the clinical utility of scalable sound-AI sleep monitoring. Support (if any)
Cho et al. (Fri,) studied this question.