We propose a method for computing joint acoustic-modulation frequency feature for speaker recognition. This feature describes the amplitude modulation spectrum of each subband, and results in a single feature vector per utterance. This vector is directly used as the speaker's modulation frequency template, excluding the need for a separate training phase. The effects of analysis parameters and pattern matching are studied using the NIST 2001 corpus. When fusing the proposed feature with the baseline MFCC/GMM system, EER is reduced from 18.2% to 16.7%
No takes yet. Share an insight, caveat, or question.
Tomi Kinnunen (2006) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: