PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 26, 2026Expert Systems1 citationsOpen Access

On the Optimal Selection of Mel‐Frequency Cepstral Coefficients for Voice Deepfake Detection

View Full Paper
SFSergio A. Falcón‐LópezLTLlanos TobarraARAntonio Robles‐Gómez

Key Points

  • The research aims to find the minimal number of Mel-Frequency Cepstral Coefficients (MFCCs) necessary for effective voice deepfake detection.
  • Analyzed the ASVspoof 2019 Logical Access dataset.
  • Utilized five traditional machine learning algorithms including Random Forest and SVM.
  • Implemented five deep learning models like CNN and ResNet.
  • Tested various numbers of MFCCs to determine optimal number for detection.
  • Assessed performance based on model accuracy and computational cost.
  • Deep learning models achieved peak performance with fewer MFCCs compared to traditional methods.
  • Traditional methods required more coefficients for stable performance, with Linear Support Vector Classification underperforming consistently.
  • Identified 32 MFCCs as effective for hybrid deployments of detection systems.

Abstract

ABSTRACT The continuous evolution of techniques for generating manipulated audio, known as voice deepfakes, and the widespread availability of tools that produce convincing forgeries have created an urgent need for reliable detection methods. This work considers the dimensionality of Mel‐Frequency Cepstral Coefficients (MFCCs) as a core design variable for practical, deployable systems. The aim is to identify the smallest number of coefficients that preserve detection performance across heterogeneous models while reducing computational cost, a critical factor for mobile and edge deployment. This study evaluates a hybrid setting on the ASVspoof 2019 Logical Access dataset, in which the same feature family serves as input to five traditional machine learning algorithms (Random Forest, k‐Nearest Neighbours, Linear Support Vector Classification, Extreme Gradient Boosting and Support Vector Machine with radial basis function kernel) and five deep learning models (Convolutional Neural Network, Recurrent Neural Network, Convolutional Recurrent Neural Network, Xception and ResNet). Results indicate that deep models reach near‐peak performance with a small number of coefficients, whereas classical methods require a larger number to achieve stable performance (except Linear Support Vector Classification, which consistently underperforms). Accordingly, 32 coefficients are considered an effective operating point for hybrid deployments. Overall, the results provide evidence to guide the selection of the number of MFCC coefficients in voice deepfake detection, aiming for efficient, reproducible and explainable systems.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Falcón‐López et al. (2026) studied this question.

synapsesocial.com/papers/69c4ccc9fdc3bde448918658https://doi.org/10.1111/exsy.70245
Ask AI
Helpful
Bookmark
Share
View Full Paper