PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 2, 2026Information1 citationsOpen Access

Morphology-Aware Segmentation and Tokenization for Turkic Languages: A CSE-Guided Framework (The Kazakh Case)

View Full Paper
UTUalsher TukeyevBRBekarys Rysbek

Key Points

  • The aim is to develop a dataset generation technology and framework for segmenting and tokenizing Turkic languages, focusing on Kazakh.
  • Developed a Complete Set of Endings (CSE) morphological model for dataset generation.
  • Proposed a CSE-guided framework for statistical tokenization and neural model segmentation.
  • Adapted tokenizers specifically for Kazakh using the new framework.
  • Incorporated Kazakh vowel-consonant harmony rules into the embedding generation.
  • Evaluated models using intrinsic, extrinsic, and external metrics on a gold-standard dataset.
  • Achieved a reduction of neural model training time by up to approximately 33%.
  • Generated the FEMSeg_kaz_v2 model trained on CSE-generated wordforms.
  • Created the FEMSeg_kaz_v3 model through training on a CSE-segmented sentence corpus.
  • Demonstrated improvements in segmentation quality through multi-level evaluation.

Abstract

The main challenge of resource-poor languages—namely, the lack of sufficiently large and linguistically informed datasets for training neural models—is addressed in this paper by developing a dataset generation technology based on a Complete Set of Endings (CSE) morphological model for Turkic languages. Building on this technology, we propose a CSE-Guided Framework for morphology-aware statistical tokenization and neural model segmentation, with Kazakh as a case study. Applying the proposed CSE-guided approach to adapt well-known tokenizers for Kazakh leads to measurable reductions in neural model training time (up to approximately 33%) in our experimental setting, primarily due to shorter tokenized sentence lengths. In addition, we extend the SOTA FEMSeg-CRF architecture by incorporating Kazakh vowel–consonant harmony rules at the embedding generation stage. Within the proposed framework, training on a corpus of CSE-generated wordforms results in the FEMSegₖazᵥ2 model, which is evaluated using intrinsic segmentation metrics. Training on a CSE-segmented sentence corpus yields FEMSegₖazᵥ3, which is further assessed using intrinsic, extrinsic, and external evaluation on a manually prepared gold-standard dataset. The paper presents a CSE-guided framework for morphology-aware tokenization and segmentation for Turkic languages, supported by corpus construction, model extensions, and multi-level evaluation. The proposed CSE-Guided Framework can potentially be adapted for other Turkic languages.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Tukeyev et al. (2026) studied this question.

synapsesocial.com/papers/6980fc17c1c9540dea80def0https://doi.org/10.3390/info17020128
Ask AI
Helpful
Bookmark
Share
View Full Paper