Speech emotion recognition has become one of the active researches in machine learning for the past few years. There are already applications that use speech emotion recognition as its feature. This paper’s purpose is to examine the difference in performance of model using multilayer perceptron (MLP), support machine vector (SVM), and Logistic Regression (LR) with Mel-frequency cepstral coefficients (MFCCs) on Indonesian language. Recording of various people’s voices are used as the dataset, which is collected using a peer-to-peer method. Emotions in the recording are classified as happy and sad. For the experiment, the authors used Precision, Recall, F1-Score, and Accuracy for the measurement to find the best model. Among three models, LR model has the perfect accuracy which is 100%. LR and MLP have the best precision rate for happy emotion and have the best recall rate for sad emotion.
No takes yet. Share an insight, caveat, or question.
Rumagit et al. (2021) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: