PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 16, 2026Systems and Soft Computing0 citationsOpen Access

End-to-end variational speech synthesis for the endangered Sylheti Nagri language

View Full Paper
MAMd. AtaullhaSJSoumik Paul JisunMRM. Shahidur Rahman

Key Points

  • This research aims to create a Text-to-Speech (TTS) system for the endangered Sylheti Nagri language to preserve its oral heritage.
  • Developed a TTS model based on the Variational Inference Text-to-Speech (VITS) framework.
  • Curated the Sylheti Nagri TTS Corpus with over 15 hours of audio from a professional voice artist, including 8268 sentences.
  • Evaluated model performance using Mean Opinion Score (MOS) and Perceptual Evaluation of Speech Quality (PESQ).
  • Achieved a Mean Opinion Score (MOS) of 3.74 (95% CI: [3.58, 3.90]), indicating good speech quality.
  • Obtained a PESQ score of 3.12 (95% CI: [2.94, 3.30]), reflecting satisfactory sound quality.

Abstract

This research work presents a novel Text-to-Speech (TTS) system for the Sylheti Nagri language, a historically significant but endangered script primarily spoken in the Sylhet region of Bangladesh. The Sylheti Nagri language remains widely spoken in daily life, but its written form in the Nagri script has been largely overshadowed by the Bangla script. To address this, we introduce the Sylheti Nagri TTS Corpus, a meticulously curated dataset comprising of over 15 h of high-quality audio recordings. This corpus, the first of its kind for Sylheti Nagri, includes 8268 sentences spoken by a professional voice artist, providing a substantial resource for preserving the language’s oral heritage. We develop an end-to-end TTS model using the Variational Inference Text-to-Speech (VITS) framework, known for its efficiency and high speech quality. Our model achieves a Mean Opinion Score (MOS) of 3.74 (95% Confidence Interval (CI): 3.58, 3.90), reflecting good overall speech quality, and a Perceptual Evaluation of Speech Quality (PESQ) score of 3.12 (95% CI: 2.94, 3.30). This work not only contributes to the preservation of the Sylheti Nagri script but also sets a foundation for future research in TTS technology for under-resourced South Asian languages.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Ataullha et al. (2026) studied this question.

synapsesocial.com/papers/6a080acea487c87a6a40cbb0https://doi.org/10.1016/j.sasc.2026.200496
Ask AI
Helpful
Bookmark
Share
View Full Paper