Randomized trial investigates voice detection in telecommunication, highlighting vulnerabilities in voice authentication systems.
The rapid progress of generative speech synthesis and voice-cloning technologies has enabled the creation of highly natural synthetic voices that pose a serious threat to telecommunication security. While most prior studies evaluate human ability to detect audio deepfakes using high-quality, studio-grade recordings, little is known about how real-world telecommunication channels affect perceptual detection. This study investigates the influence of three transmission scenarios—GSM (AMR-NB), VoLTE (AMR-WB), and VoIP with packet-loss modeling—on the human ability to distinguish natural speech from AI-generated speech. A custom speech corpus was developed, consisting of natural recordings from nine speakers and corresponding synthetic utterances generated using a state-of-the-art voice cloning system (ElevenLabs). All samples were processed through simulated telecommunication channels using real codec implementations. A listening test with 95 participants was conducted, involving binary classification (human vs. synthetic) and confidence ratings. Results show an overall detection accuracy of 54.8%, confirming that humans are poorly equipped to identify synthetic speech. Surprisingly, the highest accuracy was achieved for the narrowband GSM channel (63.7%), while VoLTE yielded the lowest performance (44.0%). The findings suggest that restricted bandwidth may emphasize prosodic irregularities typical of generative models, whereas high-quality channels mask synthetic artifacts, increasing susceptibility to voice spoofing. The results highlight the necessity of deploying additional security mechanisms in telecommunication systems relying on voice identity verification.
No takes yet. Share an insight, caveat, or question.
Warzych et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: