Key points are not available for this paper at this time.
ABSTRACT Facial expression recognition plays a critical role in applications such as digital learning, human–computer interaction, and mental state assessment. Recent advances in sensing technologies have enabled the use of multimodal data to improve recognition robustness beyond controlled laboratory environments. This study proposes a multimodal facial expression recognition framework based on an adaptive gradient–optimized convolutional weighted recurrent network (AG + CWRN). Image, audio, and sensor indicators are independently preprocessed to reduce noise and enhance discriminative patterns. Convolutional Neural Networks (CNN) are hired to extract spatial features, while an attention‐weighted recurrent architecture models temporal dependencies and performs feature‐level multimodal fusion. Adaptive Gradient (AG) optimization is applied to improve convergence and parameter stability. Experimental evaluation demonstrates that the proposed AG + CWRN framework achieves superior performance in terms of accuracy of 93.63%, precision of 89.61%, recall of 89.43%, and an F1‐score of 87.25% compared with conventional deep learning and machine learning models. The results indicate that combining multimodal feature learning with attention‐based temporal modeling delivers an actual result for facial expression recognition in digital learning environments.
Yang et al. (Mon,) studied this question.