Validation study demonstrates improved sound-based sleep quality estimation in adults, suggesting the viability of non-intrusive nighttime monitoring.
Key Points
To develop and evaluate a multimodal knowledge distillation framework that leverages comprehensive clinical sleep metrics during training to enhance sound-only subjective sleep quality prediction at inference.
Trained a multimodal teacher model on polysomnography, sleep-stage, subjective, acoustic, and demographic features using hierarchical gated variable selection networks.
Distilled the teacher model's outputs into a lightweight student model restricted solely to audio features and demographic factors.
Evaluated performance using 198 nights of recorded data from 101 adults across night-wise and subject-wise validation protocols.
Distillation improved student model prediction accuracy over non-distilled baselines, achieving the largest performance gains on participants unseen during training.
Optimal distillation parameters, learned modality weights, and feature ablation effects varied substantially between night-wise and subject-wise evaluation protocols.
A modality's standalone contribution to the teacher model did not consistently correlate with its downstream impact on distilled student performance.