Analyzing sentiment across multiple modes poses a complex challenge, requiring efficient strategies for semantic interaction and feature fusion across modalities. In this paper, we proposed a text-image sentiment classification model that utilizes multi-layer semantic enhancement and hybrid fusion strategy. First, we employed a dual-channel architecture to extract both global and local features from each modality. The self-attention mechanism was then applied to learn the internal associations within the unimodal global features for emotion classification. Concurrently, we utilized the cross-attention mechanism to thoroughly explore the semantic associations between image and text data. Subsequently, the Bidirectional Recurrent Attention Unit (BiRAU) was integrated with the self-attention mechanism to facilitate in-depth feature-level fusion, culminating in emotion prediction. Ultimately, a decision-level fusion of image, text, and multimodal classification results was executed using a dynamic weighting scheme. Experimental results on the TumEmo and MVSA-Single datasets indicate that our model enhances the performance over other related methods.
Li et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: