Computational evaluation demonstrates 85.5% phone transcription accuracy in non-native English speech, highlighting an automated framework for pronunciation training corpora.
In Computer-Assisted Pronunciation Training (CAPT), accurate phonetic transcriptions are essential for identifying mispronunciations in non-native speech corpora. Publicly available corpora for Korean English L2 learners often lack these transcriptions due to their limited size and availability. To address the shortage of accurate phonetic transcriptions, we propose a method that combines multiple Self-Supervised Learning (SSL)-based phone recognition systems with Recognizer Output Voting Error Reduction (ROVER). We trained SSL-based phone recognizers (Data2vec, Hubert, Wav2vec) on the Librispeech and CommonVoice datasets and used them to decode the L2arctic corpus. By applying ROVER, we achieved 85.5% accuracy in phone transcription compared to manual tagging. Additionally, an error analysis of 140 beginner-level sentences from the Korean Spoken English Corpus (NIA144) identified common pronunciation errors among Korean English speakers.
No takes yet. Share an insight, caveat, or question.
Kim et al. (2024) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: