PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 10, 2026Machine Vision and Applications5 citationsOpen Access

Cross-dataset video deepfake detection using Transformer and CNN architectures

GPGeorgios Petmezas

Key Points

  • The aim is to evaluate Transformer and CNN architectures for effective deepfake detection across various datasets.
  • Utilized multiple benchmark datasets and a novel facial-reenactment dataset.
  • Investigated impacts of pretraining with limited fine-tuning on smaller target subsets.
  • Analyzed the effect of temporal window length on detection performance.
  • TimeSformer achieved the highest accuracy of 78.4%, AUC of 0.801, and F1-score of 77.0%.
  • Models demonstrated improved performance with moderate fine-tuning, especially beyond 20%.
  • Longer video clips significantly enhanced performance for models designed to consider time.

Abstract

Abstract The growing sophistication of deepfake generation techniques poses serious challenges to the authenticity of digital media, with potential risks spanning privacy, security and misinformation. Deep learning (DL) methods have shown significant promise in detecting such manipulations; however, inconsistencies in application, the absence of standardized pipelines and limited cross-dataset generalization hinder their reliable deployment in real-world scenarios. This work presents a comprehensive evaluation of Transformer- and CNN-based architectures for video deepfake detection. Multiple benchmark datasets, along with a novel facial-reenactment dataset, are used to investigate cross-dataset generalization and pretraining with limited fine-tuning on small target subsets (10–30%). Additionally, we analyze the impact of temporal window length on detection performance. Experimental results demonstrate that TimeSformer consistently achieves the highest performance, reaching 78.4% accuracy, 0.801 area under the curve (AUC) and 77.0% F1-score with 96-frame clips and 30% fine-tuning, confirming the advantage of joint spatiotemporal modeling. All models benefit from moderate fine-tuning, with gains plateauing beyond 20%. Increasing clip length enhances performance for temporally aware models, highlighting the importance of extended temporal context. Overall, this study provides empirical evidence into the strengths and limitations of current architectures, offering guidance for future research and practical deployment of robust and generalizable deepfake detectors.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Georgios Petmezas (2026) studied this question.

synapsesocial.com/papers/69d895206c1944d70ce0619dhttps://doi.org/10.1007/s00138-026-01809-w
Ask AI
Helpful
Bookmark
Share
View Full Paper