Emotions, clarity, and communication habits play a key role in professional interactions. Traditional interview methods rely on subjective opinions, making it difficult for candidates to receive actionable feedback. This paper presents a multimodal AI system that analyzes recorded or live interview sessions using speech, facial expressions, and text responses to provide comprehensive feedback on interview performance. The system employs advanced AI models to transcribe speech, analyze vocal emotions and confidence, detect filler words and pauses, and evaluate facial expressions, eye contact, and emotional consistency. A late fusion strategy combines audio, text, and video features to generate confidence, clarity, and behavioral scores. Results demonstrate high accuracy in detecting confident and happy emotional states, and multimodal fusion outperforms any single modality approach. The platform also supports AI-driven mock interviews, making it a complete virtual interview preparation tool.
Gurav et al. (Wed,) studied this question.