Benchmarking ASR accuracy and conducting failure analysis on caregiver-infant interactions indicates Whisper’s limitations.
Transcribing naturalistic caregiver-child interactions is a labor-intensive task in developmental research. While automatic speech recognition (ASR) models offer potential solutions, current ASR systems, including OpenAI’s Whisper, were primarily trained on adult-directed speech (ADS) and have not been systematically evaluated for infant-directed speech (IDS). IDS differs from ADS in prosody, phonetics, utterance structure, and lexical patterns, posing unique challenges for ASR transcription. In this study, we benchmarked Whisper’s transcription accuracy on a dataset of naturalistic caregiver infant recordings, comparing ASR-generated transcripts to human-annotated gold-standard transcriptions. We evaluated Whisper’s performance using word error rate (WER) and accuracy, and conducted a failure analysis to identify systematic errors made by Whisper, including utterance segmentation mismatches, high error rates in short and repetitive speech and inconsistent filler word handling. Results indicated that Whisper struggles with short, contextually ambiguous utterances and fails to reliably segment IDS utterances. Based on these results and early user experience, we propose a preliminary pipeline for integrating ASR technology with human transcription workflows to enhance the efficiency of processing large-scale naturalistic speech data in developmental research.
No takes yet. Share an insight, caveat, or question.
Tang et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: