This paper presents the design and implementation of a multimodal emotionally intelligent virtual humanoid agent for real-time human–computer interaction. The system integrates conversational AI (Rasa), sentiment analysis (VADER), real-time gesture recognition (MediaPipe and OpenCV), speech processing, and browser-based 3D avatar rendering using GLB models. A Flask-based backend coordinates multimodal processing and ensures synchronized emotional responses across text, speech, gesture, and visual modalities. The proposed system demonstrates a scalable and lightweight framework for emotionally adaptive interaction in applications such as education, healthcare, and intelligent assistance.
Deeksha Wadhwa (Sat,) studied this question.