Los puntos clave no están disponibles para este artículo en este momento.
Emotion recognition remains a challenging task despite substantial progress in machine learning and affective computing. This study examines challenges in emotion recognition through a comparative analysis of two widely used modalities: facial images and speech signals. The analysis was conducted using FER-2013 for facial emotion recognition and the TESS and RAVDESS datasets for speech emotion recognition. A MobileNetV2-based approach was applied to visual data, while speech analysis employed MFCC-based representations and both classical and deep learning models. The study combines quantitative performance evaluation with qualitative analysis of classification behavior, focusing on emotion-specific recognition difficulties and recurring error patterns across modalities. Model performance was assessed using accuracy, precision, recall, F1-score, and confusion matrices. Across the analysed datasets, overall classification accuracy ranged from approximately 73% to 96%, while class-level F1-scores ranged from 0.48 to 0.89 depending on the emotion and modality. Happiness and surprise consistently achieved the highest recognition performance, whereas neutral emotion, fear, and disgust exhibited the lowest class-level F1-scores and generated the highest numbers of misclassifications. The experimental results confirmed that happiness and surprise achieved the highest classification performance across modalities, while neutral emotion, fear, and disgust showed reduced recognition accuracy due to weak expressive cues and overlapping feature representations. These difficulties are associated with weak or ambiguous expressive signals, overlap between emotional categories, and variability in emotional expression. The comparative findings suggest that recognition challenges arise from both modality-specific limitations and the inherent properties of emotional expression. The results highlight the importance of multimodal approaches and more flexible representations for improving emotion recognition systems.
Rafał Gasz (Mon,) studied this question.