Algorithmic study demonstrates enhanced recognition accuracy across diverse intelligibility groups, indicating effective multi-level knowledge transfer.
Automatic speech recognition (ASR) for dysarthric speech faces three major challenges: (1) data scarcity, (2) variability in speech patterns, and (3) imbalanced performance across different speech intelligibility groups. To primarily address the performance imbalance, we propose a Hierarchical Curriculum Learning (HCL) framework consisting of: (a) a hierarchical architecture that trains separate ASR models for each intelligibility group to reduce cross-group interference, (b) a multi-level knowledge distillation strategy that progressively transfers knowledge from higher to lower intelligibility groups, and (c) a fusion model that integrates predictions from group-specific ASR models. To further tackle data scarcity and variability, the proposed method incorporates meta-learning, intelligibility classification, and data augmentation tailored for dysarthric speech. Unlike prior work that adopts uniform training strategies, our approach explicitly models inter-group differences and enables effective knowledge sharing. Experiments on the UASpeech and Torgo corpus show that the proposed method achieves average WERs of 19.44% and 6.12%, respectively, with consistent and statistically significant improvements across all intelligibility groups.
No takes yet. Share an insight, caveat, or question.
Chung‐Hsien Wu (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: