PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 22, 2026Applied Sciences0 citationsOpen Access

Deepfake Speech Detection Using Perceptual Pathological Features Related to Timbral Attributes and Deep Learning

View Full Paper
ACAnuwat ChaiwongyenKZKhalid ZamanKLKanglin Li

Key Points

  • The central aim is to explore how perceptual speech-pathological features can aid in detecting deepfake speech.
  • Investigation of timbral attributes related to speech pathology.
  • Combining a deep neural network with a gammatone filterbank model.
  • Evaluation of the proposed method on the ASVspoof 2019 dataset.
  • Quantitative analysis focusing on Equal Error Rate (EER).
  • Identified meaningful distinctions between genuine and synthetic speech based on timbral attributes.
  • Achieved an EER of 5.93%, outperforming baseline detection models.
  • Demonstrated that extended dimensional representation improves detection performance.

Abstract

The detection of deepfake speech has become a significant research area due to rapid advancements in generative AI for speech synthesis. These technologies pose significant security risks in applications such as biometric authentication, voice-controlled systems, and automatic speaker verification (ASV) systems. Therefore, enhancing the detection capabilities of such applications is essential to mitigate potential threats. This study investigates perceptual speech-pathological features, which are commonly used to evaluate the unnaturalness of voice disorders in clinical settings, as potential indicators for detecting deepfake speech. Specifically, the timbral attributes of hardness, depth, brightness, roughness, sharpness, warmth, boominess, and reverberation are examined. The analysis reveals that these attributes provide meaningful distinctions between genuine and synthetic speech. Furthermore, the detection performance is enhanced by extending the dimensional representation of timbral attributes, enabling a more comprehensive characterization of the speech signal. This paper proposes a method that combines two models: one utilizing the different dimensions of speech-pathological features with a deep neural network (DNN), and another employing a gammatone filterbank model that simulates the auditory processing mechanism of the human cochlea with ResNet-18 architecture, improving deepfake speech detection. The proposed method is evaluated on the Automatic Speaker Verification Spoofing and Countermeasures Challenge (ASVspoof) 2019 dataset. Experimental results demonstrate that the proposed approach outperforms baseline models in terms of Equal Error Rate (EER), achieving an EER of 5.93%.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Chaiwongyen et al. (2026) studied this question.

synapsesocial.com/papers/699a9d8e482488d673cd3850https://doi.org/10.3390/app16042077
Ask AI
Helpful
Bookmark
Share
View Full Paper