PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 9, 2026Egyptian Informatics Journal0 citationsOpen Access

Alzheimer’s detection using audio through deep learning techniques

View Full Paper
KKKarthika KuppusamyARAnjana Rajesh

Key Points

  • This research aims to develop a robust model for early detection of Alzheimer's disease using speech analysis with deep learning techniques.
  • Developed a hybrid model called TempoBoostNet to analyze speech signals.
  • Collected a dataset of 547 speech recordings from online sources, including both patients and healthy controls.
  • Implemented features extraction techniques including Wav2Vec2.0 and Perceptual Linear Prediction for improved representation.
  • TempoBoostNet achieved 97.4% accuracy in detecting Alzheimer's disease.
  • The model also showed 97.4% precision, 97.3% recall, and 97.3% F1-score, outperforming traditional models.

Abstract

Alzheimer’s Disease (AD) is a long-lasting neurodegenerative disorder that progressively weakens memory, communication, and cognitive abilities. Early detection and diagnosis are critical for timely intervention. Recently, Artificial Intelligence (AI) models with speech analysis have emerged in detecting AD using acoustic and prosodic characteristics of individuals. However, they still achieve lower accuracy due to a lack of temporal dynamics and complex nonlinear correlation features among the speech signals. Hence, a new robust, noninvasive, and cost-effective model is proposed in this article for early detection of AD from speech signals. A new hybrid model called TempoBoostNet is developed to simultaneously capture both temporal dynamics and complex nonlinear relationships among speech signals. First, a dataset of 547 speech recordings (247 CE and 300 Healthy Controls (HC)) is gathered from publicly available Kaggle and GitHub sources. The raw audio files are preprocessed using the Wiener filter and Mahalanobis distance to eliminate noise and silence, respectively. Then, various features that reflect both low-level acoustic characteristics and higher-level speech dynamics are extracted. Also, contextual embeddings from Wav2Vec2.0 and Perceptual Linear Prediction (PLP) coefficients are extracted. All extracted features are concatenated to form a unified multi-level feature set for better feature representation. Moreover, this feature set is utilized to train the TempoBoostNet, which is built by hybridizing the Bidirectional Long Short-Term Memory and Extreme Gradient Boosting (BiLSTM-XGBoost) classifier for AD detection. Finally, experimental results show that this TempoBoostNet achieves 97.4% accuracy, 97.4% precision, 97.3% recall, and 97.3% F1-score, outperforming traditional models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kuppusamy et al. (2026) studied this question.

synapsesocial.com/papers/69fecf16b9154b0b82876207https://doi.org/10.1016/j.eij.2026.100978
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Hierarchical Age and Gender Classification from Speech Using Deep Feature Fusion and Enhanced Dimensionality Projection2025 · 1 citations
  2. 2Preclinical Alzheimer's disease: Definition, natural history, and diagnostic criteria2016 · 1,937 citations
  3. 3Dementia Detection from Speech Using Machine Learning and Deep Learning Architectures2022 · 102 citations
  4. 4Parkinson disease prediction using machine learning-based features from speech signal2023 · 43 citations
  5. 5A speech based diagnostic method for Alzheimer disease using machine learning2023 · 9 citations