PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 16, 2026Phonetics and Speech Sciences0 citationsOpen Access

Testing prosodic boundary-induced phonetic categorization in AI-based automatic speech recognition*

JJJiyoung JangRHRichard Hatcher

Key Points

  • The aim is to investigate whether prosodic boundary effects shape phonetic categorization in AI-based automatic speech recognition systems.
  • Manipulated voice onset time (VOT) of stop consonants along a voiced−voiceless continuum
  • Varying the presence of a major prosodic boundary before the target
  • Presented stimuli to a state-of-the-art ASR model (Whisper)
  • Analyzed transcription outputs for variations in voicing categorization based on prosodic context.
  • Longer VOTs led to higher probabilities of voiceless responses
  • Prosodic boundary conditions interacted with VOT
  • Whisper lacked human-like boundary-dependent shifts in voicing categories
  • Global acoustic properties influenced categorization more than VOT cues.

Abstract

Phonetic categorization is shaped by systematic relationships between segmental cues and prosodic structure. In human speech perception, stop consonant voicing varies according to prosodic boundary strength, reflecting boundary-conditioned patterns of phonetic realization. This study examines whether such prosodic effects are observable in AI-based automatic speech recognition (ASR). We manipulated the voice onset time (VOT) of word-initial stop consonants along an English voiced−voiceless continuum while varying the presence of a major prosodic boundary preceding the target. The stimuli were presented to a state-of-the-art ASR model (Whisper), and the transcription outputs were analyzed to determine how voicing categorization varied across prosodic boundary contexts. Results showed an effect of VOT, with longer VOTs yielding higher probabilities of voiceless responses. While prosodic boundary condition interacted with VOT, Whisper did not exhibit a human-like boundary-dependent shift in the voicing category boundary, and these effects were contingent on the voicing of the original token and place of articulation. Particularly, global acoustic properties associated with the source exerted stronger influence on categorization, often overriding VOT cues. These findings suggest that while Whisper encodes sufficient acoustic detail to support coarse phonetic categorization, it does not recalibrate segmental cue interpretation for prosodic boundary structure. This study highlights a fundamental divergence between human perceptual normalization and end-to-end ASR inference, with implications for prosody-sensitive modeling of speech perception and recognition.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jang et al. (2026) studied this question.

synapsesocial.com/papers/69e07dc72f7e8953b7cbeb82https://doi.org/10.13064/ksss.2026.18.1.065
Ask AI
Helpful
Bookmark
Share
View Full Paper