Robust text-independent speaker identification using Gaussian mixture speaker models

Key Points

Key points are not available for this paper at this time.

Abstract

This paper introduces and motivates the use of Gaussian mixture models (GMM) for robust text-independent speaker identification. The individual Gaussian components of a GMM are shown to represent some general speaker-dependent spectral shapes that are effective for modeling speaker identity. The focus of this work is on applications which require high identification rates using short utterance from unconstrained conversational speech and robustness to degradations produced by transmission over a telephone channel. A complete experimental evaluation of the Gaussian mixture speaker model is conducted on a 49 speaker, conversational telephone speech database. The experiments examine algorithmic issues (initialization, variance limiting, model order selection), spectral variability robustness techniques, large population performance, and comparisons to other speaker modeling techniques (uni-modal Gaussian, VQ codebook, tied Gaussian mixture, and radial basis functions). The Gaussian mixture speaker model attains 96.8% identification accuracy using 5 second clean speech utterances and 80.8% accuracy using 15 second telephone speech utterances with a 49 speaker population and is shown to outperform the other speaker modeling techniques on an identical 16 speaker telephone speech task.>

Connected Papers

Building similarity graph...

Analyzing shared references across papers

Discussion

Authors

D.A. Reynolds

MIT Lincoln Laboratory

Richard C. Rose

University of Strathclyde

Journals

IEEE Transactions on Speech and Audio Processing

Actions

Institutions

AT&T (United States)

MIT Lincoln Laboratory

Systems Technology (United States)

References and Citations

Connected Papers

Building similarity graph...

Analyzing shared references across papers

Discussion

Cite this study

Reynolds et al. (Sun,) studied this question.

synapsesocial.com/papers/6a079b45b2d9a7d54307aabc — DOI: https://doi.org/10.1109/89.365379

Also consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

Digital Communications· 1984 · 1,386 citations
Long-term feature averaging for speaker recognition· 1977 · 83 citations
Vector Quantization· 1984 · 2,385 citations
Effectiveness of linear prediction characteristics of the speech wave for automatic speaker identification and verification· 1974 · 923 citations
Text-independent talker identification with neural networks· 1991 · 67 citations

Also consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

Digital Communications· 1984 · 1,386 citations
Long-term feature averaging for speaker recognition· 1977 · 83 citations
Vector Quantization· 1984 · 2,385 citations
Effectiveness of linear prediction characteristics of the speech wave for automatic speaker identification and verification· 1974 · 923 citations
Text-independent talker identification with neural networks· 1991 · 67 citations

Robust text-independent speaker identification using Gaussian mixture speaker models

Key Points

Abstract

Citation Network

Connected Papers

Discussion

Authors

Journals

Actions

Institutions

References and Citations

Citation Network

Connected Papers

Discussion

Cite this study

Also consider

Also consider