PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 15, 2026IEICE Transactions on Information and Systems0 citationsOpen Access

End-to-End Spontaneous Speech Recognition Based on Disfluency Labeling

KHKoharu HoriiMFMeiko FukudaKOKengo OHTA

Key Points

  • The aim is to enhance automatic speech recognition by addressing errors caused by disfluent utterances.
  • Proposed disfluency labeling to categorize segments of speech as fillers or hesitations.
  • End-to-end training of the ASR model incorporating labeled disfluency data.
  • Evaluation experiments comparing the new method against previous ASR techniques.
  • The disfluency labeling method achieved significantly higher recognition accuracy.
  • Explicit learning of disfluency features as labels improved the model’s ability to interpret spontaneous speech.

Abstract

Disfluent utterances in spontaneous speech, such as fillers and hesitations, cause recognition errors during automatic speech recognition (ASR). To address this problem, we propose a method called “disfluency labeling”, which replaces disfluent segments in transcription data with one of two labels: # (filler) or @ (hesitation). End-to-end training of the ASR model with such labeled data enables recognition of these disfluent segments as targets, like characters, allowing the extraction of what the speaker intended to say. In evaluation experiments, the proposed disfluency labeling method achieved higher recognition accuracy than the previous proposed ASR methods treating disfluencies, suggesting that explicit learning of disfluency features as labels is effective for improving spontaneous speech recognition.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Horii et al. (2026) studied this question.

synapsesocial.com/papers/69b64c67b42794e3e660da8fhttps://doi.org/10.1587/transinf.2025edp7157
Ask AI
Helpful
Bookmark
Share
View Full Paper