PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 12, 2025Information2 citationsOpen Access

Speech Recognition and Synthesis Models and Platforms for the Kazakh Language

View Full Paper
AKAidana KaribayevaVKVladislav KaryukinBABalzhan Abduali

Key Points

  • Strongest performance for speech-to-text was achieved by Soyle (BLEU 74.93, WER 18.61), indicating effective language adaptation.
  • KazakhTTS2 delivered the most natural quality in text-to-speech with DNSMOS 8.79–8.96, suggesting high perceptual quality.
  • The study assessed various models like GPT-4 Transcribe and Whisper, revealing significant differences in accuracy and performance.
  • Technical barriers include the Kazakh language's agglutinative structure and rich vowel system, affecting model effectiveness.

Abstract

With the rapid development of artificial intelligence and machine learning technologies, automatic speech recognition (ASR) and text-to-speech (TTS) have become key components of the digital transformation of society. The Kazakh language, as a representative of the Turkic language family, remains a low-resource language with limited audio corpora, language models, and high-quality speech synthesis systems. This study provides a comprehensive analysis of existing speech recognition and synthesis models, emphasizing their applicability and adaptation to the Kazakh language. Special attention is given to linguistic and technical barriers, including the agglutinative structure, rich vowel system, and phonemic variability. Both open-source and commercial solutions were evaluated, including Whisper, GPT-4 Transcribe, ElevenLabs, OpenAI TTS, Voiser, KazakhTTS2, and TurkicTTS. Speech recognition systems were assessed using BLEU, WER, TER, chrF, and COMET, while speech synthesis was evaluated with MCD, PESQ, STOI, and DNSMOS, thus covering both lexical–semantic and acoustic–perceptual characteristics. The results demonstrate that, for speech-to-text (STT), the strongest performance was achieved by Soyle on domain-specific data (BLEU 74.93, WER 18.61), while Voiser showed balanced accuracy (WER 40.65–37.11, chrF 80.88–84.51) and GPT-4 Transcribe achieved robust semantic preservation (COMET up to 1.02). In contrast, Whisper performed weakest (WER 77.10, BLEU 13.22), requiring further adaptation for Kazakh. For text-to-speech (TTS), KazakhTTS2 delivered the most natural perceptual quality (DNSMOS 8.79–8.96), while OpenAI TTS achieved the best spectral accuracy (MCD 123.44–117.11, PESQ 1.14). TurkicTTS offered reliable intelligibility (STOI 0.15, PESQ 1.16), and ElevenLabs produced natural but less spectrally accurate speech.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Karibayeva et al. (2025) studied this question.

synapsesocial.com/papers/68ebffcfdef9fcb308ff23c3https://doi.org/10.3390/info16100879
Ask AI
Helpful
Bookmark
Share
View Full Paper