Objectives: This study aimed to examine how two features—sentence pitch contour and word position—affect speech perception in noise among children and whether these effects differ between musicians and nonmusicians. Design: The study included 51 children (20 musicians and 31 nonmusicians) aged 93 to 179 months (7.75 to 14.92 years) with normal hearing. Participants completed word identification tasks using the Sung Speech Corpus, a set of closed-set matrix sentences. Stimuli included sentences of five words varying in three sentence pitch contours—naturally spoken, mixed-pitch (different pitch for each word), and fixed-pitch (same pitch for each word)—at two signal-to-noise ratios (SNRs) (0, +3 dB) using speech-shaped steady-state noise based on all the words combined. Neurocognitive functions, including nonverbal IQ, receptive vocabulary, and auditory short-term memory (STM) capacity, were assessed. Linear mixed-effects models analyzed the main effects and two-way interactions of pitch contour, word position in the sentence, SNR, and musicianship on word identification accuracy. Models were conducted with and without controlling for neurocognitive factors correlated with word identification accuracy to assess their impact. Recency-to-primacy differences (RP-differences) were calculated, and a separate linear mixed-effects model was used to examine the effects of pitch contour, SNR, and musicianship on RP-differences. Results: Significant main effects of pitch contour and SNR, as well as their interaction, were observed for word identification accuracy. Performance improved progressively from fixed-pitch to mixed-pitch to naturally spoken contours and with increasing SNR. The interaction indicated that unnatural pitch contours exacerbated the decline in word identification accuracy under the lower SNR. A U-shaped pattern emerged for word identification as a function of word position, with lower performance in the middle of a sentence and highest performance at the initial or final position, reflecting primacy and recency effects, respectively. The primacy effect was more negatively affected by unnatural pitch contours (mixed-pitch and fixed-pitch) than the recency effect, as indicated by larger RP-differences for distorted contours compared with the spoken contour. Musicians and nonmusicians had comparable nonverbal IQ and receptive vocabulary, but musicians demonstrated higher auditory STM capacity. While musicians outperformed nonmusicians in overall word identification, this advantage became nonsignificant when auditory STM capacity was controlled for. No significant interactions were found between musicianship and sentence pitch contour or SNR, suggesting that musician advantage did not specifically counteract pitch contour distortions or unfavorable SNRs. Both groups exhibited comparable RP-differences, indicating that enhanced auditory STM did not provide additional benefits for recalling words in challenging positions. Conclusions: These findings suggest that both types of unnatural pitch contours—mixed-pitch and fixed-pitch—reduced children’s speech perception in noise compared with the natural spoken contour, with the fixed-pitch condition producing the poorest performance. The primacy and recency effects in free recall were evident in school-aged children’s word identification with five-word sequences, with primacy more susceptible to pitch contour distortions. Musicians’ higher speech perception in noise was primarily attributed to their greater auditory STM capacity. However, this advantage did not extend to resilience against pitch contour distortions, unfavorable SNRs, or challenging word positions for recall.
Firdose et al. (Tue,) studied this question.