This paper presents the development of a compact, near-real-time anger detection system through voice analysis, implemented on an ESP32 microcontroller. The system is designed to address domestic violence scenarios in which anger often precedes aggression. Due to the computational constraints of embedded devices, we propose an optimized Convolutional Neural Network based on VGG16, called MiniVGG16, which reduces the number of parameters while maintaining an 80% accuracy in anger recognition. The system processes audio signals in near real-time, extracting Mel-Frequency Cepstral Coefficients from speech data. These features are then fed into the MiniVGG16 model, which has been optimized using TensorFlow Lite Micro for efficient inference on the ESP32. The entire pipeline, including audio sampling, feature extraction, and classification, is executed in 676.4 ms, enabling practical deployment in near-real-world applications. The system’s small memory usage and real-time processing capability make it a cost-effective and accessible tool for emotion monitoring in home environment, mental healthcare scenarios, and security applications. Future improvements include noise filtering and data augmentation to improve classification performance under diverse acoustic conditions.
No takes yet. Share an insight, caveat, or question.
Fernandez-Morales et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: