PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 23, 2006IEEE Transactions on Audio Speech and Language Processing271 citations

Enriching speech recognition with automatic detection of sentence boundaries and disfluencies

View Full Paper
YLYang LiuESE. ShribergASAndreas Stolcke

Key Points

Key points are not available for this paper at this time.

Abstract

Effective human and automatic processing of speech requires recovery of more than just the words. It also involves recovering phenomena such as sentence boundaries, filler words, and disfluencies, referred to as structural metadata. We describe a metadata detection system that combines information from different types of textual knowledge sources with information from a prosodic classifier. We investigate maximum entropy and conditional random field models, as well as the predominant hidden Markov model (HMM) approach, and find that discriminative models generally outperform generative models. We report system performance on both broadcast news and conversational telephone speech tasks, illustrating significant performance differences across tasks and as a function of recognizer performance. The results represent the state of the art, as assessed in the NIST RT-04F evaluation

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Liu et al. (2006) studied this question.

synapsesocial.com/papers/6a206e6b3d50bdc5d1029b8dhttps://doi.org/10.1109/tasl.2006.878255
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A corpus-based study of repair cues in spontaneous speech1994 · 210 citations
  2. 2Modeling dynamic prosodic variation for speaker verification1998 · 141 citations
  3. 3Repeating Words in Spontaneous Speech1998 · 445 citations
  4. 4Comparing and Combining Generative and Posterior Probability Models: Some Advances in Sentence Boundary Detection in Speech2004 · 35 citations