PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 4, 2026Scientific Reports0 citationsOpen Access

Multi-emotion and intensity-driven response generation for richer multimodal dialogue

ASApoorva SinghRSRaj ShreeDPDhirendra Pandey

Key Points

  • The aim is to generate responses in dialogue that accurately reflect multiple emotions and their intensity, enhancing interactions with AI systems.
  • Developed a Multimodal Multi Emotion and Intensity-leaded Deliberation Decoder (MMEI-DD) framework.
  • Used a Transformer network to encode multimodal knowledge from text, audio, and visual elements.
  • Implemented a two-step decoding mechanism to generate responses based on specified emotions and intensity.
  • Demonstrated that the MMEI-DD framework improves response generation compared to existing methods.
  • Showed enhanced performance and reliability in integrating multimodal knowledge for dialogue responses.

Abstract

The combination of well-chosen words becomes a thought, although the thought attains meaningful significance only when spoken with emotion. The significance of every sentence could vary depending on the emotions behind it. As AI speech agents evolve, understanding and integrating human emotions is key to fostering natural and impactful interactions. Recently, researchers have given preference to the development of communication systems that are tuned to emotions. In recent work, our focus is on generating multi-emotional responses with precisely controlled intensity by integrating essential insights from text, audio, and visual elements in multimodal dialogues. In this regard, the authors designed a Multimodal Multi Emotion and Intensity-leaded Deliberation Decoder (MMEI-DD) framework that encodes the multimodal knowledge using the Transformer network and captures the inter-modality representation. The desired emotion and corresponding intensity are provided during the two-step decoding mechanism to the deliberation decoder to generate multi-emotion and intensity guided responses. The authors expanded the multimodal feature set (MEIMD + +) based on the recently introduced MEIMD (Multiple Emotion Intensity-Aware Multi-party Dialogue) dataset and demonstrate that integrating multimodal knowledge enhances the performance and reliability of our proposed approach compared to the existing method. In this paper our proposed model leverage to support further research, the necessary resources and code will be made publicly available.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Singh et al. (2026) studied this question.

synapsesocial.com/papers/69d0afb4659487ece0fa5b50https://doi.org/10.1038/s41598-026-41034-z
Ask AI
Helpful
Bookmark
Share
View Full Paper