Cohort study reveals temporal and spectral speech markers track depression across independent sites, indicating acoustic analysis is a scalable diagnostic aid.
Key Points
To assess the cross-site reproducibility, robustness, and generalizability of objective acoustic speech markers for detecting major depressive disorder in an independent clinical cohort.
Analyzed acoustic data from two independent psychiatric cohorts in Germany comprising 135 participants (71 healthy controls and 64 patients with major depressive disorder) performing positive and negative storytelling tasks.
Extracted over 80 temporal, lexical, and spectral acoustic features, correlating them with Beck Depression Inventory (BDI-II) scores.
Trained machine learning models on the RWTH Aachen cohort and evaluated cross-site classification performance on the University of Oldenburg cohort.
Temporal and spectral features, such as utterance duration, pause duration, and MFCCs, consistently associated with depression diagnosis and symptom severity across both cohorts.
Machine learning models trained on one cohort achieved an area under the receiver operating characteristic curve (ROC-AUC) of 0.63 when evaluated on the independent cohort.
Voice quality metrics including shimmer and jitter showed site-dependent inconsistencies during negative storytelling, indicating sensitivity to recording environments rather than stable pathology.