Computational evaluation demonstrates improved emotion classification on multi-party dialogue benchmarks, indicating that selective gating and frozen encoders effectively mitigate multimodal noise.
This paper addresses multimodal emotion recognition in multi-party dialogue, combining transcripts, speech waveforms, and video of speakers’ faces. This article presents Multimodal Attention Fusion Transformer (MAFT), a framework built around various persistent difficulties such as: neutral utterances account for close to half of all labels, the transcript channel outperforms audio and video by a wide margin, and the non-text signals on this particular dataset suffer from laugh-track contamination, frequent camera switches, and occluded faces. MAFT presents an attention-based multimodal fusion framework evaluated on the MELD benchmark using frozen pretrained encoders. The current fusion mechanism pairs every combination of two modalities through cross-attention, then applies a per-pair sigmoid gate an explicit kill switch that lets the network zero out a channel when it degrades the prediction for a given input. Modality importance starts at a hand-picked ratio of 0.5/0.3/0.2 (text/audio/video); a lightweight MLP conditioned on conversational context adjusts these on every sample. Training uses cosine-annealed class-level curriculum scheduling that rotates emphasis from high-frequency emotions toward rare ones over 50 epochs, alongside a combined focal and supervised contrastive objective. All feature encoders are frozen: RoBERTa-large (1024-d), HuBERT-large (1024-d), MARLIN (768-d). The system reaches 69.2% weighted F1 on MELD’s test partition on this earliest prototype, which relied on base-scale encoders without any of these mechanisms, managed 57.3%.
No takes yet. Share an insight, caveat, or question.
Rajagopal et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: