Randomized trial demonstrates improved hate speech detection in multimodal memes, indicating potential for better moderation.
Memes act as cryptic tools for sharing sensitive ideas, often requiring contextual knowledge to interpret. It makes multimodal memes moderation a difficult task, as existing works either lack high-quality datasets on nuanced hate categories or rely on low-quality social media visuals. Here, we curate two novel multimodal hate speech datasets comprising English memes – MHS and MHS-Con, which capture fine-grained hateful abstractions in regular and confounding scenarios, respectively. We benchmark these datasets against several competing baselines. Furthermore, we introduce SAFE-MEME (Structured reAsoning FramEwork), a novel multimodal Chain-of-Thought-based framework employing Q&A-style reasoning (SAFE-MEME -QA) and hierarchical categorization (SAFE-MEME -H) to enable robust hate speech detection in memes. SAFE-MEME -QA outperforms the strongest open-source baseline model, showing an improvement of \(2\%\) on MHS and \(4.7\%\) on MHS-Con and closely follows the closed-source models, GTP-4o and Gemini 2.5. In contrast, SAFE-MEME -H surpasses SAFE-MEME -QA and shows an improvement of \(3\%\) over the best open-source baseline or equivalent performance to GPT-4o only on MHS. We show that fine-tuning a single-layer adapter within SAFE-MEME-H outperforms fully fine-tuned models in regular fine-grained hateful meme detection. However, the fully fine-tuning approach with a Q&A setup is more effective for handling confounding cases. We also systematically examine the error cases, offering valuable insights into the robustness and limitations of the proposed structured reasoning framework for analyzing hateful memes. 1
No takes yet. Share an insight, caveat, or question.
Nandi et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: