In this paper, we use several techniques with conventional vocal feature extraction (MFCC, STFT), along with deep-learning approaches such as CNN, and also context-level analysis, by providing the textual data, and combining different approaches for improved emotion-level classification. We explore models that have not been tested to gauge the difference in performance and accuracy. We apply hyperparameter sweeps and data augmentation to improve performance. Finally, we see if a real-time approach is feasible, and can be readily integrated into existing systems.
No takes yet. Share an insight, caveat, or question.
Huang et al. (2019) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: