Randomized trial demonstrates improved Parkinsonian speech classification in patients using auditory models, suggesting better diagnosis methods.
This study aims to build a machine learning model to differentiate Parkinsonian speech (PS) from normal speech using models that reflect the transformation of the sound in the auditory system (auditory models). Parkinson’s is known to influence speech articulation for 89% of the patients. The impacted speech is described as breathier and rougher compared to that of healthy individuals. Previous studies have predicted PS phonation using mel-frequency cepstral coefficients (MFCCs), Short-Time Fourier Transform and acoustic measures, such as F0, its jitter and shimmer. In this study, features related to the perception of speech using outputs from auditory models are explored. Since several phonation types can be perceived under the same category, despite production differences, our approach could potentially improve the accuracy of automatic PS classification. We compared predictions made with MFCCs, the output from a Gammatone filterbank, and from the Cascade of Asymmetric Resonators with Fast-Acting Compression (CARFAC) model (including its basilar membrane, inner hair cell potential, and neural activity patterns (NAPs) stages). We also compared predictions made with Bidirectional Long Short-Term Memory (BiLSTM) and Transformer architectures. Best results so far have been obtained with a BiLSTM model trained on a NAPs output, with a classification accuracy of 80.19%.
No takes yet. Share an insight, caveat, or question.
Sakai et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: