PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 11, 2026Journal of Intelligent Information Systems3 citationsOpen Access

Multimodal misinformation detection across diverse languages using RAG and LLMs

SHSheetal HarrisVTVinh Thong TaMTMarcello Trovati

Key Points

  • The research aims to improve the detection of multimodal fake news across diverse languages using advanced AI frameworks.
  • Developed a Multilingual & Multimodal Retrieval-Augmented Generation framework (M&M-RAG).
  • Utilized Large Vision-Language Models (LVLMs) and Large Language Models (LLMs) for news verification.
  • Integrated real-time multilingual evidence retrieval and cross-modal reasoning for fact-checking.
  • Established a comprehensive dataset for multimodal fake news detection in Urdu.
  • Achieved state-of-the-art performance with 94.6% accuracy and 94.2% F1 score.
  • Outperformed existing models like SpotFake and MPFN in multimodal fake news detection.
  • Demonstrated robustness in zero-shot and cross-lingual scenarios without fine-tuning.

Abstract

The rapid spread of multimodal fake news (FN) on Online Social Networks (OSNs) threatens digital information ecosystems, particularly in low-resource languages. Existing multimodal fake news detection (FND) methods are largely limited to high-resource settings, restricting their global applicability. We propose an M&M-RAG, a Multilingual & Multimodal Retrieval-Augmented Generation framework, that leverages Large Vision-Language Models (LVLMs) and Large Language Models (LLMs) to verify news claims across English, Chinese and Urdu. M&M-RAG integrates real-time multilingual evidence retrieval, language-aware prompting, and cross-modal reasoning for fact verification. We further propose Multi-Ax-to-Grind Urdu, the first large-scale, multi-domain multimodal benchmark for FND in Urdu. Experiments on typologically diverse monolingual multimodal datasets demonstrate that M&M-RAG achieves state-of-the-art (SOTA) performance, with 94.6% accuracy and 94.2% F1 score, surpassing models such as SpotFake, MPFN, MMCFND, and Semi-FND. The proposed framework remains robust in zero-shot and cross-lingual scenarios under frozen-model inference without task-specific fine-tuning. The results underscore the scalability and interpretability of LVLM-based approaches for combating multimodal misinformation, particularly in under-represented and typologically diverse languages.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Harris et al. (2026) studied this question.

synapsesocial.com/papers/69d9e57078050d08c1b75ac1https://doi.org/10.1007/s10844-026-01042-x
Ask AI
Helpful
Bookmark
Share
View Full Paper