PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 12, 202432 citationsOpen Access

Careless Whisper: Speech-to-Text Hallucination Harms

View Full Paper
AKAllison KoeneckeACAnna Seo Gyeong ChoiKMK. Mei

Key Points

Key points are not available for this paper at this time.

Abstract

Speech-to-text services aim to transcribe input audio as accurately as possible. They increasingly play a role in everyday life, for example in personal voice assistants or in customer-company interactions. We evaluate Open AI's Whisper, a state-of-the-art automated speech recognition service outperforming industry competitors, as of 2023. While many of Whisper's transcriptions were highly accurate, we find that roughly 1\% of audio transcriptions contained entire hallucinated phrases or sentences which did not exist in any form in the underlying audio. We thematically analyze the Whisper-hallucinated content, finding that 38\% of hallucinations include explicit harms such as perpetuating violence, making up inaccurate associations, or implying false authority. We then study why hallucinations occur by observing the disparities in hallucination rates between speakers with aphasia (who have a lowered ability to express themselves using speech and voice) and a control group. We find that hallucinations disproportionately occur for individuals who speak with longer shares of non-vocal durations -- a common symptom of aphasia. We call on industry practitioners to ameliorate these language-model-based hallucinations in Whisper, and to raise awareness of potential biases amplified by hallucinations in downstream applications of speech-to-text models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Koenecke et al. (2024) studied this question.

synapsesocial.com/papers/68e796dbb6db643587707adchttps://doi.org/10.1145/3630106.3658996
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Automatic Speech Recognition System to Record Progress Notes in a Mobile EHR: A Pilot Study2024 · 2 citations
  2. 2Quantifying and Improving the Performance of Speech Recognition Systems on Dysphonic Speech2023 · 9 citations
  3. 3Survey of Hallucination in Natural Language Generation2022 · 4,352 citations
  4. 4The Role of ChatGPT, Generative Language Models, and Artificial Intelligence in Medical Education: A Conversation With ChatGPT and a Call for Papers2023 · 878 citations
  5. 5Analyzing Qualitative Data2007 · 4,077 citations