This work presents an innovative approach to test the Arabic language proficiency assessment via Automatic Speech Recognition (ASR) by enhancing the proficiency of the Whisper model in transcribing Arabic speech. The core of our research involved fine-tuning the Whisper model using a substantial, large-scale Arabic speech corpus, with a specific focus on Modern Standard Arabic. This process used a 2000-h Arabic-labeled speech corpus, the QASR dataset, and improved the model’s Word Error Rate (WER). After optimization, the fine-tuned Whisper model’s WER was reduced from 35% to 7% on the QASR dataset, corresponding to an absolute reduction of 28 percentage points (approximately 80% relative reduction). These results demonstrate the strong generalization ability of the fine-tuned model across multiple Arabic ASR benchmarks. A key component of our methodology was the development of a sophisticated scoring system. This system integrates various similarity metrics, such as cosine similarity, the Jaccard index, and the Levenshtein distance, with a machine learning regression model. This multifaceted system provides a comprehensive assessment of reading proficiency, proposing a practical automated assessment method that contributes to the field of AI language transcription and to its application in the assessment of students’ reading. Our research also introduces the ICONET dataset, an augmented Arabic speech corpus comprising 3160 h of diverse and tailored audio–text pairs designed for fine-tuning ASR models. This study demonstrates the potential of fine-tuning pretrained models for specific linguistic contexts (Arabic), establishing a foundation for future research in ASR and language technology.
Badawi et al. (Sat,) studied this question.