Los puntos clave no están disponibles para este artículo en este momento.
This work is concerned with devising a robust Parkinson's (PD) disease detector from speech in real-world operating conditions using (i) foundational models, and (ii) speech enhancement (SE) methods. To this end, we first fine-tune several foundational-based models on the standard PC-GITA (s-PC-GITA) clean data. Our results demonstrate superior performance to previously proposed models. Second, we assess the generalization capability of the PD models on the extended PC-GITA (e-PC-GITA) recordings, collected in real-world operative conditions, and observe a severe drop in performance moving from ideal to real-world conditions. Third, we align training and testing conditions applaying off-the-shelf SE techniques on e-PC-GITA, and a significant boost in performance is observed only for the foundational-based models. Finally, combining the two best foundational-based models trained on s-PC-GITA, namely WavLM Base and Hubert Base, yielded top performance on the enhanced e-PC-GITA.
Building similarity graph...
Analyzing shared references across papers
Loading...
Quatra et al. (Sun,) studied this question.
synapsesocial.com/papers/68e63ae4b6db6435875cc76e — DOI: https://doi.org/10.48550/arxiv.2406.16128
Moreno La Quatra
Università degli Studi di Enna Kore
Maria Francesca Turco
Torbjørn Svendsen
Norwegian University of Science and Technology
Building similarity graph...
Analyzing shared references across papers
Loading...