PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 8, 20250 citationsOpen Access

Second language Korean Universal Dependency treebank v1.2: Focus on data augmentation and annotation scheme refinement

View Full Paper
HSHakyung SungGSGyu‐Ho Shin

Key Points

  • Fine-tuning language models on well-annotated L2 datasets significantly enhances their performance.
  • The revised annotation guidelines align closely with the Universal Dependencies framework for better consistency.
  • This study includes 5,454 additional sentences to strengthen the L2 Korean language model training.
  • Results indicate a marked performance improvement across various metrics for in-domain and out-of-domain datasets.

Abstract

We expand the second language (L2) Korean Universal Dependencies (UD) treebank with 5,454 manually annotated sentences. The annotation guidelines are also revised to better align with the UD framework. Using this enhanced treebank, we fine-tune three Korean language models and evaluate their performance on in-domain and out-of-domain L2-Korean datasets. The results show that fine-tuning significantly improves their performance across various metrics, thus highlighting the importance of using well-tailored L2 datasets for fine-tuning first-language-based, general-purpose language models for the morphosyntactic analysis of L2 data.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sung et al. (2025) studied this question.

synapsesocial.com/papers/68e62de1a8c0c6d45873fea4https://doi.org/10.48550/arxiv.2503.14718
Ask AI
Helpful
Bookmark
Share
View Full Paper