Multimodal Sentiment Analysis (MSA) represents an advancing research domain focused on identifying sentiment through Audio(A), Video (V), and Text (T) modalities. A significant challenge lies in capturing joint representations by associating information across various modalities with integration techniques. Numerous existing approaches rely on acquiring joint representation through the concatenation of input features without fully exploiting interactions to ensure consistency with complementarity among modalities. To address this problem, a novel framework named Joint Representation with Optimized Transformer (JRT) has been developed to overcome the challenges in multimodal sentiment analysis through hierarchical interactions among modalities. The process begins with a diverse dataset comprising text, audio, and visual data from the CMUMOSI and CMU-MOSEI datasets. Preprocessing steps are performed with normalization, noise reduction, and feature scaling to ensure data quality. Feature extraction is used to isolate meaningful patterns from each modality for further analysis. The framework integrates these features into a unified representation through a Joint Representation Translator (JRT), which aligns heterogeneous data for compatibility using cyclic translation. This translation process captures joint representations of bimodality by translating one modality into another forward with backward passes using encoderdecoders to maintain consistency between modalities. To explore complementarity among modalities, a transformer-based prediction mechanism is optimized with the Adaptive Dragon Optimization Algorithm (ADOA), which strengthens unimodal features with common information extracted from bimodality, enhancing model accuracy and convergence. Extensive experiments conducted with CMU-MOSI and CMU-MOSEI datasets validate the effectiveness of this framework, showing superior performance compared to existing methods. The proposed method achieves 95.11% accuracy (%), 97.51% recall (%), 95.14% F-Score (%), 0.90 correlation coefficient(unitless), 0.051 false positive rate (%), 97% negative predictive value (%), and 0.0489 mean squared error (MSE). These results highlight substantial advancements in accuracy and robustness over traditional approaches.
Vasanthi et al. (Wed,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: