Randomized trial demonstrates effective speaker recognition from disguised voices, indicating strong implications for forensic applications.
This research work focuses on a speaker identification framework that effectively recognizes speakers in both normal and disguised voice conditions. Disguised speech presents considerable difficulties in forensic contexts because the normal voice (NV) is physically changed with voice masking on the mouth (MM), objects in the mouth (OM), pinched nostrils (PN), and whispering (WP). Here, Audacity version 3.7.7 was used to record speech samples in a controlled environment. A hybrid Mel-frequency cepstrum coefficient and long short-term memory (MFCC-LSTM) deep network model is proposed to address this issue for identification of speakers and computation of their performance parameters. It uses both advanced auditory feature extraction techniques and deep temporal learning techniques. For feature extraction, MFCC, delta MFCC (ΔMFCC), and double delta MFCC (ΔΔMFCC) coefficients along with related statistical coefficients, such as mean and correlation coefficients, are used to compute the features. The qualities consist of both the voice signal's fixed spectrum characteristics and its dynamic temporal characteristics. After feature extraction, the LSTM classifier was used to compute the performance parameter by keeping long-term frame dependencies in patterns, which solves the difficulties with classical machine learning techniques. Here computed accuracy, precision, recall, specificity, and F1-score to verify how well the proposed system worked. The results indicate that the proposed MFCC-LSTM method achieved classification accuracies of 90.67%, 89.48%, 90.17%, and 91.00% for WP, PN, MM, and OM, respectively, demonstrating a strong ability to produce effective classification results. The proposed MFCC-LSTM model demonstrates a significant improvement associated with the existing MFCC-SVM method. The accuracy result was increased from 52.40% to 57.80%, precision increased from 49.20% to 52.18%, recall increased from 48.50% to 57.80%, and the F1-score increased from 50.45% to 54.75%. Additionally, specificity increases significantly from 48.70% to 51.86%, representing that the proposed technique attains distinctly improved classification performance. The proposed method, the MFCC-LSTM approach, increases the overall accuracy from 52.40% to 57.80%. It also yields high per-disguise accuracies of 90.67% (WP), 89.48% (PN), 90.17% (MM), and 91.00% (OM). The performance parameter shows that combining cepstral feature extraction with deep learning types makes it considerably easier to detect disguised speakers. This has large consequences for forensic speech investigation and biometric security classifications.
No takes yet. Share an insight, caveat, or question.
Mahesh K. Singh (2026) studied this question.
Synapse has enriched 2 closely related papers on similar clinical questions. Consider them for comparative context: