PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 16, 20260 citationsOpen Access

A CTC-Aligned Lightweight Korean TTS Pipeline without External Aligners

View Full Paper
MJMinhyung Jo

Key Points

  • The research aims to develop a Korean text-to-speech (TTS) system that operates without requiring external alignment tools.
  • Develop a CTC aligner for aligning mel-spectrogram frames to phonemes.
  • Organize training into two stages: jointly train the CTC aligner with the acoustic model and then retrain the acoustic model using fixed alignments.
  • Utilize a lightweight HiFi-GAN for vocoding.
  • Achieved a Mean Opinion Score (MOS) of 3.67 using the proposed method on the KSS dataset.
  • Outperformed the MFA-based TTS approach, which scored 3.46 at the same model scale.

Abstract

Lightweight text-to-speech (TTS) models have gained attention for their fast inference and low resource requirements, but they typically depend on external aligners such as the Montreal Forced Aligner (MFA) during training. For Korean, the publicly available MFA acoustic models and pronunciation dictionaries are of limited quality, requiring substantial manual effort to set up a reliable alignment pipeline. In this paper, we propose a lightweight Korean TTS pipeline that can be trained without any external aligner. The proposed method introduces a Connectionist Temporal Classification (CTC) aligner that classifies mel-spectrogram frames directly into phonemes and extracts phoneme durations via forced alignment. Training is organized into two stages: in Stage 1, the CTC aligner is jointly trained with the acoustic model to obtain alignments; in Stage 2, the acoustic model is retrained from scratch using the fixed alignment targets extracted by the Stage 1 aligner. While the monotonic alignment search (MAS) used by larger models fails to converge in lightweight configurations because the encoder representation lacks sufficient capacity, CTC alignment operates directly on the mel spectrogram and is therefore robust to the encoder's capacity. For vocoding, we use a lightweight HiFi-GAN fine-tuned on Korean speech data. On the KSS dataset, the proposed method achieves a MOS of 3.67 without any external aligner, outperforming an MFA-based counterpart (3.46) at the same model scale.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Minhyung Jo (2026) studied this question.

synapsesocial.com/papers/69e07e242f7e8953b7cbf1c7https://doi.org/10.5281/zenodo.19564646
Ask AI
Helpful
Bookmark
Share
View Full Paper