Disfluent utterances in spontaneous speech, such as fillers and hesitations, cause recognition errors during automatic speech recognition (ASR). To address this problem, we propose a method called “disfluency labeling”, which replaces disfluent segments in transcription data with one of two labels: # (filler) or @ (hesitation). End-to-end training of the ASR model with such labeled data enables recognition of these disfluent segments as targets, like characters, allowing the extraction of what the speaker intended to say. In evaluation experiments, the proposed disfluency labeling method achieved higher recognition accuracy than the previous proposed ASR methods treating disfluencies, suggesting that explicit learning of disfluency features as labels is effective for improving spontaneous speech recognition.
Horii et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: