Abstract—The rapid advancement of deepfake technology poses significant threats to information authenticity, identity protection, and societal trust. This paper presents a survey of multimodal deepfake detection frameworks with a particular emphasis on combining Efficient Temporal Modeling for Classification (ETMC) in video analysis and RawNet-based audio analysis. By merging temporal-spatial and acoustic cues, such frameworks achieve strong accuracy while remaining computationally practical. The survey explores unimodal and multimodal detection methods, discusses their strengths and limitations, and highlights the need for robust, lightweight, and real-time detection mechanisms for safeguarding digital media integrity. Index Terms—Deepfake detection, multimodal framework, video forensics, audio analysis, media security, Synthetic media, Deep learning, Forgery detection, Information integrity
A et al. (Sat,) studied this question.