The increasing realism of deepfake videos has intensified the need for reliable video-level detection systems, but benchmark performance must be interpreted together with generalization, temporal-modeling, and efficiency limits. This study presents a reproducible multi-backbone framework for detecting manipulated videos on the official Celeb-DF v2 benchmark. Five pretrained image-classification architectures—ResNet-50, EfficientNet-B4, ConvNeXt-Small, ViT-Base, and Swin-Base—are fine-tuned under a unified protocol using uniform frame sampling, class-balanced training, RandAugment, MixUp, CutMix, random erasing, label smoothing, AdamW optimization, cosine learning-rate scheduling, and exponential moving average weights. During inference, frame-level fake probabilities are stabilized using horizontal-flip test-time augmentation and aggregated into video-level predictions by mean probability pooling. This aggregation is a fixed probability-pooling rule rather than an explicit temporal model. A probability-level ensemble of ResNet-50, ConvNeXt-Small, and Swin-Base combines convolutional and attention-based representations. On the official 518-video Celeb-DF v2 test set, the top-three ensemble achieves a video-level AUC of 99.967% and an average precision of 99.983%, while ResNet-50 provides the strongest single-model accuracy–efficiency trade-off. Additional analyses examine frame-to-video aggregation, ROC and precision–recall behavior, probability distributions, frame-probability stability over sampled frames, a proof-of-concept temporal-splice sensitivity test, and inference efficiency. The results demonstrate highly competitive in-dataset performance on Celeb-DF v2 only. Because no external benchmark testing, learned temporal baseline, confidence-gated cascade, or broad partial-manipulation benchmark is included, the results should not be interpreted as evidence of cross-dataset robustness, in-the-wild deployment readiness, learned temporal reasoning, or general localization capability. Cross-dataset evaluation on FaceForensics++, DFDC, WildDeepfake, and related benchmarks, probability calibration, false-positive control, explicit temporal modeling, cascade-based inference, and expanded localization evaluation are identified as priority future-work directions.
Alshalfi et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: