PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 5, 2026Scientific Reports0 citationsOpen Access

A dual-domain spatio-temporal and frequency framework for robust deepfake detection

ASArman SajjadiSSSayna SarvarMNMobin Nekou

Key Points

  • This study aims to enhance deepfake detection by leveraging a dual-domain framework that incorporates spatio-temporal and frequency cues.
  • Introduced DSTF-Net, comprising two coordinated branches: WaveFormer for frequency representation and SFormer for spatial analysis.
  • WaveFormer applies 2D Haar wavelet decomposition and transformer-based temporal combiner for multi-scale insights.
  • SFormer uses Swin-Transformer and temporal encoder to capture local and long-range motion dynamics within 32-frame sequences.
  • Accuracy varied from 50.41% to 62.21% depending on training-testing direction and manipulation category.
  • AUC values ranged from 0.4776 to 0.7613.
  • Performance declines highlight challenges with generalization across datasets and compression types.

Abstract

The rapid proliferation of high-fidelity facial synthesis has made automated deepfake detection a core requirement in digital forensics. Existing detectors that rely exclusively on spatial or temporal cues can be sensitive to compression artifacts and variations among manipulation pipelines. This paper introduces DSTF-Net, a dual-domain spatio-temporal and temporal-frequency framework designed to combine complementary evidence from these two domains. DSTF-Net consists of two coordinated branches. The WaveFormer branch applies a three-level 2D Haar wavelet decomposition to face-aligned frames, generating multi-scale frequency representations that expose seams, resampling traces, and texture inconsistencies, and then aggregates these frame-level descriptors using a transformer-based temporal combiner with hybrid statistical pooling. In parallel, the SFormer branch utilizes a Swin-Transformer backbone combined with a temporal encoder to jointly model spatial structure and motion dynamics across sequences of 32 frames, capturing both local appearance distortions and long-range temporal irregularities. The 512-dimensional embeddings from the two branches are fused using an adapted MLP-Mixer that performs token- and channel-wise mixing to learn a compact, discriminative video-level representation. However, cross-dataset evaluation without target-domain fine-tuning reveals a substantial performance decline, with accuracy ranging from 50.41% to 62.21% and AUC ranging from 0.4776 to 0.7613, depending on the training–testing direction and manipulation category. These results indicate that generalization to unseen dataset and compression characteristics remains challenging. Overall, the results demonstrate the effectiveness of integrating spatio-temporal and frequency-domain cues for intra-dataset deepfake detection, while highlighting the need for additional domain-generalization and compression-invariant learning strategies before deployment in unconstrained environments.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Sajjadi et al. (2026) studied this question.

synapsesocial.com/papers/6a72e79c226790f370656e50https://doi.org/10.1038/s41598-026-65079-2
Ask AI
Helpful
Bookmark
Share
View Full Paper