Key points are not available for this paper at this time.
To address the limitations of conventional pronunciation assessment in handling acoustic variability. This research presents a novel framework for Japanese pronunciation assessment that integrates self-supervised speech representations with temporal alignment to facilitate granular feedback. The proposed methodology utilizes Wav2Vec 2.0 for automated, high-precision word segmentation, followed by dynamic time warping (DTW) to quantify similarity in pitch-accent patterns. Experimental results indicate that the long short-term memory (LSTM)-based classification model achieves an accuracy of 92.5% with an F1-score of 0.92, demonstrating high reliability in pronunciation discrimination. Furthermore, the system effectively isolates prosodic deviations through word-level distance heatmaps, providing actionable diagnostic feedback for learners. This study contributes a robust, model-driven pipeline that enhances the diagnostic capability of computer-assisted pronunciation training (CAPT) systems for Japanese language learning.
Bundasak et al. (Thu,) studied this question.