Speech Emotion Recognition (SER) is a key enabling technology for advanced human–computer interaction and affective computing. This paper presents an adaptive hybrid SER framework that combines a deep neural feature extraction module with a heterogeneous ensemble of machine learning classifiers, including XGBoost, Support Vector Machines (SVMs), and Random Forest. To overcome the limitations of static fusion strategies, a confidence-gated meta-classification mechanism is introduced to dynamically weight the contribution of each base classifier according to its instance-level reliability. The proposed approach is evaluated on two widely adopted benchmark datasets, EmoDB and SAVEE, achieving competitive accuracies of 98.88% and 91.92%, respectively. Experimental results demonstrate that the proposed fusion strategy significantly improves robustness against inter-speaker variability and emotional ambiguity, while maintaining low computational complexity suitable for real-time implementation. These findings highlight the effectiveness of the proposed framework as a robust and efficient solution for speech emotion recognition. While the model is evaluated on benchmark datasets, it is intended as a foundational component for future emotion-aware systems, including applications in human–computer interaction.
Titouni et al. (2026) studied this question.