Purpose: This study examined the extent to which spectral degradation and speaking style (child-directed vs. adult-directed speech) affect emotion recognition from prosodic cues and how these effects are modulated by concurrent tasks involving nonauditory sensory input. Method: Adults with normal hearing completed an emotion recognition task under three conditions: alone (auditory single-task), concurrently with a low-load visual memory task (four identical images), and with a high-load visual memory task (four different images). Stimuli consisted of semantically neutral sentences spoken in five emotions (angry, happy, neutral, sad, and scared) and two speaking styles (child-directed and adult-directed). All sentences were vocoded to simulate spectral degradation. Emotion recognition was assessed using a single-interval, five-alternative, forced-choice paradigm, in which the participants were asked to indicate which of five emotions was associated with each heard sentence. Results: Emotion recognition was significantly reduced for vocoded stimuli, as indicated by lower sensitivity ( d ') and prolonged reaction times (RTs). Child-directed speech led to better performance than adult-directed speech, although its facilitative effect was reduced under vocoded conditions. Dual-tasking impaired performance, with lower d ' values in both dual-task conditions and slower RTs under high-load dual-task conditions. Crucially, dual-task effects did not significantly vary with spectral degradation or speaking style. Conclusions: Top-down cognitive demands from cross-modal dual-tasking and bottom-up stimulus factors, such as spectral degradation and speaking style, independently influence emotion recognition from prosodic cues. These findings provide insight into how cochlear implant users perceive emotional speech in complex, multimodal environments.
Zilong Xie (Thu,) studied this question.