Comparative study reveals early fusion strategies outperform late fusion in emotion recognition systems.
Multimodal emotion recognition systems typically rely on one of two fusion strategies: feature-level (early) fusion, which learns joint audio-visual representations before classification, or decision-level (late) fusion, which combines independently trained unimodal classifiers afterward. Controlled comparisons between the two strategies under matched experimental conditions remain rare in the literature. This paper presents AffectFusion, a comparative framework that evaluates both strategies under two conditions. The first is a cross-dataset setting combining CREMA-D audio with FER2013 facial imagery; the second is a within-dataset setting using synchronized RAVDESS audio-visual recordings. Two feature-level architectures were developed: a Gated Feature Fusion network for the cross-dataset condition, and a Multimodal Self-Attention Fusion network for the within-dataset condition. These were evaluated alongside a decision-level Smart Hybrid Fusion mechanism that combines entropy-based reliability weighting, confidence comparison, modality expertise weighting, and cross-modal agreement boosting. Twenty architectural variants were benchmarked during development before the best-performing models were evaluated independently on held-out test data. The cross-dataset Gated Feature Fusion model reached 66.88 percent accuracy, exceeding the cross-dataset Smart Hybrid Fusion model's 64.80 percent; both clearly outperformed unimodal audio and visual baselines of 57.20 percent and 57.40 percent. The within-dataset self-attention model reached 55.66 percent accuracy, substantially outperforming simpler gated alternatives evaluated under the same synchronized-modality condition. Within the experimental settings investigated here, these results indicate that feature-level fusion holds a modest but consistent advantage over decision-level fusion under synthetically paired cross-dataset conditions, and that self-attention-based fusion is most beneficial when genuine audio-visual synchronization is available for it to exploit. A carefully engineered decision-level combination rule can nonetheless narrow the gap with feature-level fusion substantially, offering a practical alternative where joint feature access across modalities is unavailable.
No takes yet. Share an insight, caveat, or question.
Himanshu Dixit (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: