Randomized trial develops synsets for Uzbek language, indicating a novel method for low-resource languages.
The rapid expansion of digital information has intensified demand for intelligent language processing systems for morphologically rich, low-resource languages. This paper presents the linguistic foundations of an electronic thesaurus for Uzbek, employing a seven-stage hybrid methodology that combines multilingual BERT (mBERT) and Word2Vec embeddings with expert linguistic validation to construct synsets encoding synonymic, antonymic, hyperonymic, hyponymic, meronymic and troponymic relations. Drawing on UzNatCorpora and ARANEUM UZBEKIUM, and validated against the five-volume Uzbek Explanatory Dictionary (O’TIL) – an extensively documented resource comprising over 80,000 entries compiled by Uzbek lexicographers on the basis of written and oral corpora, whose role in the methodology is discussed in Section 3.2 – we develop 67 pilot synsets across four grammatical categories. The lexemes were selected using a stratified criterion described in Section 3.2. Query-expansion accuracy (86%) was measured against a human-annotated relevance judgement set of 200 queries (inter-annotator κ = 0.81) using McNemar’s test (χ2 = 7.84, p = 0.005, two-tailed, contingency table in Table 3); synonym coverage (88%) was assessed against the Uzbek Synonyms Dictionary (Mahmudov et al.). The hybrid method substantially outperforms both purely manual and purely automatic approaches (p < 0.05). The conceptual model addresses Uzbek-specific challenges including polysemy, homonymy and ideographic field formation, providing a replicable framework for WordNet-type thesaurus construction in under-resourced Turkic languages.
No takes yet. Share an insight, caveat, or question.
Muhammadjon Najmiddinov (2026) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: