Key points are not available for this paper at this time.
Sarcasm often arises from subtle contrasts between literal meaning and speaker intention. As online communication increasingly includes voice-based content, detecting sarcasm across speech and text becomes more important—and more complex. The existing methods usually focus on generic multimodal fusion but often miss how sarcasm manifests differently in each modality. We propose a model that explicitly encodes audio signals into the textual representation space, allowing prosodic cues to inform language understanding. To extract relevant features at different levels, we use a multi-scale convolutional architecture. The experiments show consistent gains over prior models on both text and speech sarcasm detection tasks.
Wu et al. (2025) studied this question.