PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 18, 20260 citationsOpen Access

Deep Vision in Detecting Visual Deception

View Full Paper
ITInternational Journal for Research In Science & Advanced Technologies

Key Points

  • The research aims to develop a multi-modal deep learning model for identifying manipulated visual and audio content.
  • Developed Deep Vision AI combining video, image, and audio analysis.
  • Utilized models like Xception, LSTM, and EfficientNetB0 for assessments.
  • Merged benchmark datasets to enhance model strength and generalization.
  • Implemented a web-based application for real-time media detection.
  • Achieved 92% accuracy in video analysis, 86% in image analysis, and 87% in audio.
  • Total system accuracy reached 90% using a majority voting fusion mechanism.
  • Multi-modal integration improved detection reliability compared to single-modality methods.

Abstract

The intensive development of deep learning made it possible to produce highly realistic manipulated content, such as face-swapped videos, forged images, and synthetic audio, which are very dangerous in the form of misinformation, identity fraud, and other criminal activities. This paper introduces Deep Vision AI, a multi-modal deep learning model of identifying manipulated media with the main emphasis on the production of fake content. The suggested system combines video, image, and audio analysis with such sophisticated models as Xception with Long Short-Term Memory (LSTM) to model the video sequence, EfficientNetB0 to analyze images in the context of video forensics, and MFCC-based to extract features and classify audio with the help of the ASVspoof dataset. Several benchmark datasets, such as FaceForensics++, Celeb-DF and DeepFake Detection datasets are merged to improve generalization and strength. The results of the experiment show that the proposed system has an accuracy of 92, 86 and 87 percent on video, image and audio respectively and has a total system accuracy of 90 percent with a majority voting fusion mechanism. The system is deployed as a web-based application through the use of Flask allowing real-time identification of manipulated media. The findings are that multi-modes integration greatly enhances reliability of detection relative to single-modes.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

International Journal for Research In Science & Advanced Technologies (2026) studied this question.

synapsesocial.com/papers/69e320e740886becb65400f0https://doi.org/10.5281/zenodo.19604704
Ask AI
Helpful
Bookmark
Share
View Full Paper