Los puntos clave no están disponibles para este artículo en este momento.
This paper presents an introduction to various deep learning techniques with the aim of capturing and classifying emotional state from speech utterances. Architectures such as Convolutional Neural Network(CNN) and Long Short-Term Memory(LSTM) have been used to test the emotion capturing capability from various standard speech represenations such as mel spectrogram, magnitude spectrogram and Mel-Frequency Cepstral Coefficients (MFCC's) on two popular datasets- EMO-DB and IEMOCAP. Experimental findings along with reasoning have been presented as to which architecture and feature combination is better suited for the purpose of speech emotion recognition. This work explores the widely used basic deep learning architectures used in literature.
Pandey et al. (Mon,) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: