PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 29, 2026Computers, materials & continua/Computers, materials & continua (Print)0 citationsOpen Access

A Prosody-Guided Multi-Stream Framework for Universal Detection of AI-Synthesized Speech across Codec and Vocoder Domains

View Full Paper
AAAkmalbek AbdusalomovMMMukhriddin MukhiddinovFAFakhriddin Abdirazakov

Key Points

  • To develop a detection model that effectively identifies AI-synthesized speech across various synthesis methods.
  • Proposed UniTector++, a multi-stream detection architecture with prosody awareness.
  • Utilized Whisper-based semantic embeddings and high-level prosodic features.
  • Incorporated Multi-Domain Adaptive Graph Attention Fusion for stream integration.
  • Achieved state-of-the-art performance with an average Equal Error Rate of 0.57%.
  • Outperformed competitive baselines by 28% in various unseen synthesis scenarios.

Abstract

Recent advancements in AI-synthesized speech have resulted in highly realistic deepfake audio, posing severe threats to authentication systems and digital media trust. Existing detection models struggle to generalize across diverse synthesis methods, especially those involving neural codec-based Audio Language Models (ALMs). In this work, we propose UniTector++, a novel prosody-aware, multi-stream detection architecture that generalizes across vocoder- and codec-based synthesis. UniTector++ incorporates three complementary streams—Whisper-based semantic embeddings, high-level prosodic features, and codec artifact representations—fused through a Multi-Domain Adaptive Graph Attention Fusion (MAGAF) module. Furthermore, an Emotion-Consistency Verification Module (ECVM) reinforces alignment between speech style and prosodic content, and a Universal Adversarial Robustness (UAR) head improves resistance against adversarial attacks. Evaluated on three benchmark datasets—ASVspoof2021, PolyFake, and Codecfake—UniTector++ achieves state-of-the-art performance with average Equal Error Rate (EER) of 0.57% under unseen synthesis scenarios, outperforming competitive baselines by a relative margin of 28%. Our results demonstrate the model’s superior generalization, interpretability, and robustness, offering a significant advancement in universal deepfake speech detection.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Abdusalomov et al. (2026) studied this question.

synapsesocial.com/papers/69f154e0879cb923c49451b1https://doi.org/10.32604/cmc.2026.080444
Ask AI
Helpful
Bookmark
Share
View Full Paper