Stuttering and related speech disorders can interrupt the natural flow of speech through repetitions, prolonged sounds, and pauses or delays, affecting millions of people worldwide. Although considerable progress has been made in artificial intelligence-based Automatic Speech Recognition (ASR) technology, most of the current models remain mainly designed for high resource and dominant lingua franca languages, e.g., English, and underperform for Arabic. This paper presents insights into stuttering classification in Arabic using Whisper ASR from OpenAI, trained to classify speech segments as fluent or disfluent. As a necessary first step, we frame the task as a binary (fluent vs. disfluent) classification. Fine-grained recognition of specific disfluency types (such as repetitions, prolongations, and blocks) is left to future work. The most salient aspect of this study is the construction of a marked stuttering speech corpus in which speech segments of fluent and disfluent speech were collected from real clinical cases. A systematic comparative framework is built between the full Whisper family (Tiny–Large) and the Wav2Vec2.0 family (Base–XLarge) under identical conditions, providing the first benchmark for Arabic stuttering classification. We find that Whisper outperforms Wav2Vec2.0 at every scale, including the smallest variants, and remains reliable even for low-resource deployment. This confirms that the Whisper encoder is suitable for clinical Arabic speech-disorder workflows.
Alnamazi et al. (Thu,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: