This paper presents an extensive empirical study aiming to identify the optimal combination of feature extraction techniques and machine learning algorithms, including deep learning, for automated mispronunciation detection during Quran recitation. Three feature extraction methods-Mel-Frequency Cepstral Coefficients (MFCC), Spectrogram, and Kaldi pitch - are systematically evaluated in conjunction with a diverse set of classifiers, including Gradient Boosting (GB), Support Vector Machines (SVM), Logistic Regression (LR), Random Forest (RF), Artificial Neural Network (ANN), Convolutional Neural Networks (CNN), and Long Short-Term Memory (LSTM) networks. The study leverages the QDAT dataset, comprising 1500 audio samples from 150 readers, with metadata specifying age, gender, and three distinct rules for mispronunciation detection. Experimental results demonstrate that the integration of MFCC with deep learning classifiers consistently achieves the highest performance across multiple evaluation metrics. Model robustness is validated using receiver operator characteristic curves (ROC) and accuracy metrics.
No takes yet. Share an insight, caveat, or question.
Alsahafi et al. (2024) studied this question.