PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 18, 2025IEEE Transactions on Neural Systems and Rehabilitation Engineering2 citationsOpen Access

Convolution-Augmented Transformers for Enhanced Speaker-Independent Dysarthric Speech Recognition

View Full Paper
ZZZihan ZhongQWQianli WangSSSatwinder Singh

Key Points

  • The proposed dysarthric speaker-independent models improved word recognition accuracy by 21.9% for isolated speech.
  • Using a Conformer-based system, the models achieved significant word error reduction of 18.5% for continuous speech.
  • Cross-dataset validation was introduced, showing the models' limitations in handling severe dysarthria transcription.
  • Pre-training on standard speech data enabled effective adaptation to two dysarthric datasets.

Abstract

Dysarthria is a motor speech disorder characterized by muscle movement difficulties that complicate verbal communication. It poses significant challenges to Automatic Speech Recognition (ASR) systems due to data scarcity and speaker variability among dysarthric individuals. This study investigates speaker-independent (SI) approaches to assist speakers with communication impairments. Firstly, we developed dysarthric SI models using a Conformer-based system and a three-stage transferlearning pipeline that employs a selective layer freezing PEFT strategy to mitigate data scarcity. We pre-trained on standard speech and progressively adapted the models to two dysarthric datasets, respectively. Secondly, we introduced a benchmark framework for evaluating the generalizability of SI models with cross-dataset validation-a previously unexplored approach in dysarthric ASR, providing a more realistic scenario. The results demonstrate that the proposed dysarthric SI models outperform all baseline systems. Specifically, on the TORGO dataset, our models improved word recognition accuracy by 21.9% for isolated speech and reduced the word error rate by 18.5% for continuous speech. On UA-Speech, our optimal dysarthric SI model achieved a word recognition improvement of 14.6% over Whisper and 28.3% over the base model for isolated speech. Nevertheless, our cross-dataset testing showed that models tended to produce isolated words when asked to transcribe continuous speech for severe dysarthria, highlighting the need to further improve SI generalization.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Zhong et al. (2025) studied this question.

synapsesocial.com/papers/68d461b631b076d99fa607b6https://doi.org/10.1109/tnsre.2025.3610792
Ask AI
Helpful
Bookmark
Share
View Full Paper