Multimodal sarcasm detection identifies ironic intent by jointly analyzing text and images. It has attracted increasing attention due to its importance in understanding user-generated content on social media. However, multimodal observations are often incomplete due to data loss during transmission or collection, leading to unreliable predictions in real-world multimodal sarcasm detection scenarios. To address incomplete observations, we further propose a Reliability-Aware Dynamic Fusion (RADF) module, which predicts the reliability of the textual, visual, and interactive views from their representations, converts these reliability scores into dynamic fusion weights via a temperature-scaled softmax to control the sharpness of the weight distribution, and refines the fused features through feature-wise scaling. In this way, degraded views are suppressed while more informative views are emphasized under incomplete observations. Extensive experiments on public datasets validate the effectiveness of our approach, which consistently outperforms existing baselines.
Guo et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: