Key points are not available for this paper at this time.
Deepfake detection has emerged as a critical research area in image content security. To address the issue of limited generalization caused by insufficient modeling of forgery representations, this paper proposes a channel-aware local–global representation learning network for generalizable deepfake detection. Specifically, we introduce a Local–Global Integration Vision Transformer (LGI-ViT) block that learns local and global representations and further integrates them with input features to capture more generalizable forgery cues at both fine-grained and global levels. Local representation learning is enhanced through coordinate convolution, while a hybrid convolution–Transformer architecture is employed to model global dependencies. Based on these representations, a residual connection is incorporated to integrate the combined local–global representation with the original input features. In addition, a Lightweight Channel Attention Network (LCAN) block is designed to strengthen interactions among feature channels and improve the discriminability of forgery-related representations. Experimental results demonstrate that the proposed network, trained on the FaceForensics++ (FF++) dataset, achieves cross-dataset AUC scores of 73.95% and 79.42% on the DeepFake Detection Challenge (DFDC) and Celeb-DF datasets, respectively. It outperforms the best-performing baseline among 11 competing models for generalizable detection by 1.01 percentage points on average, thereby validating its effectiveness in generalizable deepfake detection.
Kuang et al. (Mon,) studied this question.