Comparative evaluation demonstrates improved dialect identification in older male Korean speakers using wav2vec 2.0 XLS-R, highlighting the value of linguistically informed categories.
Previous Korean automatic dialect identification systems have demonstrated limited accuracy, hovering around 60%, and have often relied on traditional machine learning models. A further methodological issue has been the categorization of dialects by administrative districts, a method inconsistent with established linguistic boundaries. This study seeks to overcome these limitations by applying a modern, Transformer-based, end-to-end speech model, wav2vec 2.0 cross-lingual speech representation (XLS-R). The model, pre-trained on approximately 436,000 hours of speech from 128 languages, was fine-tuned for this task. We also re-categorized the dialects into six major linguistic zones: Central, Gyeongnam, Gyeongbuk, Jeonnam, Jeonbuk, and Jeju. Using a dataset of middle-aged and senior male speakers, the XLS-R model was compared against baseline models like Bi-LSTM. The results show a significant performance increase, with the XLS-R model achieving an F1 score of 86.8%—a 14 percentage point improvement over the strongest baseline. While confusion between certain dialects persists, this research validates the effectiveness of applying large-scale, pre-trained models to the nuanced task of dialect identification and underscores the importance of using linguistically-informed categories. Furthermore, the findings contribute to advancing dialect identification technology for applications in speech recognition and forensic science.
No takes yet. Share an insight, caveat, or question.
Lee et al. (2025) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: