PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 23, 20250 citationsOpen Access

METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark

View Full Paper
YXYang XuQZQi ZhangSJShuming Jiang

Key Points

  • METER offers enhanced broad modality coverage for forgery detection, allowing more nuanced interpretations of results.
  • The benchmark includes four tracks requiring real-vs-fake classification and evidence-chain-based explanations for each modality.
  • METER implements a unique human-aligned training strategy that enhances the explainability of detection outcomes.
  • The unified approach improves the interpretability of AI-generated content assessments in safety-critical applications.

Abstract

With the rapid advancement of generative AI, synthetic content across images, videos, and audio has become increasingly realistic, amplifying the risk of misinformation. Existing detection approaches predominantly focus on binary classification while lacking detailed and interpretable explanations of forgeries, which limits their applicability in safety-critical scenarios. Moreover, current methods often treat each modality separately, without a unified benchmark for cross-modal forgery detection and interpretation. To address these challenges, we introduce METER, a unified, multi-modal benchmark for interpretable forgery detection spanning images, videos, audio, and audio-visual content. Our dataset comprises four tracks, each requiring not only real-vs-fake classification but also evidence-chain-based explanations, including spatio-temporal localization, textual rationales, and forgery type tracing. Compared to prior benchmarks, METER offers broader modality coverage and richer interpretability metrics such as spatial/temporal IoU, multi-class tracing, and evidence consistency. We further propose a human-aligned, three-stage Chain-of-Thought (CoT) training strategy combining SFT, DPO, and a novel GRPO stage that integrates a human-aligned evaluator with CoT reasoning. We hope METER will serve as a standardized foundation for advancing generalizable and interpretable forgery detection in the era of generative media.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Xu et al. (2025) studied this question.

synapsesocial.com/papers/68d473bb31b076d99fa6ca3fhttps://doi.org/10.48550/arxiv.2507.16206
Ask AI
Helpful
Bookmark
Share
View Full Paper