Objectives: Neural responses evoked by speech stimuli may help assess hearing in infants and young children who are unable to reliably participate in behavioral hearing tests. These responses entrain to the periodicity of the speech envelope and are therefore referred to as envelope-following responses (EFRs). EFRs occur at the voice’s fundamental frequency (f0; usually >80 Hz) and lower frequency (<10 Hz) periodicities associated with, for example, phoneme and syllable transitions. This study focuses on statistical approaches for inferring frequency-specific and non-frequency-specific audibility of speech from f0 (f0-EFRs) and slow-rate (<10 Hz) EFRs (SR-EFRs). The objective was to improve test sensitivity by combining f0-EFRs and SR-EFRs across multiple speech stimuli. Design: Test sensitivity was assessed using electroencephalogram recordings from 66 normal-hearing participants (22 adults, 44 infants) in response to a modified male-spoken /susa∫i/ (“susashee”) presented monaurally. The /susa∫i/ token evokes SR-EFRs and eight f0-EFRs: three f0-EFRs were evoked by the low frequency, first formants of the /u/, /a/, and /i/ vowels, three by the mid-frequency higher formants of the vowels, and two by the high frequency /s/ and /∫/ fricatives. The f0-EFRs were first analyzed separately, per fricative and per vowel formant, using Hotelling’s T 2 test (HT 2 ) or a modified HT 2 test (T 2 Diag ). Vowels and fricatives were subsequently pooled within each frequency band (low, mid, high), and “combination tests” were conducted to assess the audibility of each band. Two approaches were assessed: (1) all HT 2 - or T 2 Diag -generated p values were pooled, per frequency band, and evaluated while accounting for multiple comparisons using the Bonferroni or inverse χ 2 approach, or (2) a single high-dimensional f0-EFR feature set was constructed, per frequency band, and evaluated with a single HT 2 or T 2 Diag hypothesis test. Non-frequency-specific combination tests were also evaluated, which leveraged either all eight f0-EFRs simultaneously or all eight f0-EFRs in combination with SR-EFRs. Test specificity was evaluated using no-stimulus electroencephalogram background activity recorded from 10 adults. Results: For the frequency-band-specific assessments using f0-EFRs, test sensitivity was highest for the inverse χ 2 method: Compared with the non-combined (phoneme- and formant-specific) f0-EFR analyses, detection rates increased by 0.04, up to 0.35. When leveraging all eight f0-EFRs simultaneously in the non-frequency-specific audibility assessment, detection rates increased by up to 0.55 relative to the low- and mid-frequency-band assessments, but only by 0.03 compared with the high-frequency-band assessment. The non-frequency-specific assessment, leveraging f0-EFRs and SR-EFRs simultaneously, led to a small (nonsignificant) increase (up to 0.02) in detection rates compared with using f0-EFRs alone. Detection rates for the SR-EFRs were lower than those of the f0-EFRs but approached 100% detection after ~30 min of test time in some test conditions. Conclusions: Combining speech-evoked EFRs elicited by multiple phonemes and phoneme formants improved test sensitivities in frequency-band-specific and non-frequency-specific assessments of speech audibility. Such combination tests might be used to quickly obtain a preliminary, non-frequency-specific assessment of hearing, and then progressively home in on frequency-specific assessments. If testing terminates early, for example, due to the infant waking up, then the preliminary assessments may still provide useful information.
Chesnaye et al. (Tue,) studied this question.