This study investigates a two-step forced alignment method for L2 English speech produced by Korean middle school students with varying proficiency levels. Specifically, we analyzed 3254 utterances—1627 each from 13-year-old female speakers and male speakers—drawn from an educational corpus provided by AI Hub, a Korean government-funded platform providing open datasets for AI research. The middle school students of varying proficiency levels read speech materials, which often contained multiple sentences per utterance and exhibited considerable variation in sentence length and lexical difficulty. Prior studies (Rousso et al., 2024; Williams et al., 2024) report that MFA performs well for canonical utterances but is less robust to disfluencies and mismatches with transcripts. WhisperX, while flexible in disfluent conditions, is sensitive to long or complex utterances, often producing alignment drift. Following the approach of Coulange et al. (2024), we applied WhisperX for initial word-level alignment and then refined phoneme boundaries using MFA. This two-step method yielded more temporally precise and linguistically consistent alignments across both short and long utterances, particularly for spontaneous and low-proficiency speech. The improved alignment supports accurate segmental and prosodic acoustic analyses and demonstrates the value of combining ASR-based and traditional tools for phonetic studies of adolescent L2 learners.
Tae-Jin Yoon (Wed,) studied this question.