Framework shows high accuracy in emotion classification for assistive robotics, suggesting enhanced interactions for communication disabilities.
Multimodal emotion recognition has become an important function for intelligent humanrobot interaction, especially in an assistive robotics scenario for individuals with communication disabilities. The paper presents a Hybrid BERT-Vision Transformer (BERT-ViT) system constructed to combine complementary information obtained from natural language and facial expression analysis for enhanced emotional comprehension in assistive response systems. By employing BERT to produce contextual embeddings for language and the Vision Transformer to obtain fine-grained visual features, the proposed BERT-ViT-domain framework performs cross-modal emotion classification in a robust manner. Contrastive learning is applied for the multimodal feature alignment, and then the fusion step concatenates the [CLS] token embeddings from both modalities. The evaluation of the model was conducted with the MELD dataset, and the findings indicate promising performance associated with a mean accuracy on seven emotion classes and performance levels above conventional unimodal and early-fusion counterparts. These results also indicate the model to perform relatively well at detecting subtle and complex affective states of concern in relation to ASD support for empathetic interaction, speech impairments, and socially assistive robotics. The modular nature of the model's architecture provides for seamless integration into assistive platforms, while also supporting feedback and personalization features. The experimental validation includes linguistic profiling, space-time processing of emotions, and performance classification measures of accuracy (98.71%), precision (98.34%), recall (98.00%), and F1-score (98.38%). Also noteworthy is the strong resistance to overfitting with training-validation convergence and low overall error rates (FPR: 1.58%, FNR: 0.98%). This hybrid framework contributes to realizing advanced multimodal emotion recognition and developing a scalable, responsive, and privacy-responsible solution toward emotionally aware assistive robotic applications in sensitive health care and educational contexts.
No takes yet. Share an insight, caveat, or question.
Abbas et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: