PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 14, 2026The Journal of the Acoustical Society of America0 citations

Towards machine learning-driven speech intelligibility prediction models: Examining relationships with spectrotemporal modulation

View Full Paper
KYKatsuhiko Yamamoto

Key Points

  • This study aims to evaluate the effectiveness of machine learning-driven speech intelligibility prediction models that incorporate spectrotemporal modulation features.
  • Analyzed traditional auditory-based prediction models for speech intelligibility.
  • Discussed the integration of spectrotemporal modulation features in prediction models and machine learning-based approaches.
  • Highlighted trends in the Clarity Project utilizing large-scale data from listening experiments.
  • Machine learning models incorporating spectrotemporal modulation showed enhanced prediction accuracy for speech intelligibility.
  • Findings from the Clarity Project indicate significant improvements in predictive performance for hearing-impaired listeners.
  • Advanced speech foundation models demonstrate potential for practical applications in auditory diagnostics.

Abstract

This study examines the progression from auditory-based speech intelligibility prediction to models incorporating spectrotemporal modulation features, and ultimately to machine learning models, from an engineering perspective. Prediction models based on auditory characteristics are well-suited for evaluating speech enhancement technologies in hearing aids and other auditory assistive devices. Traditional models based on auditory filterbank and modulation filterbank have provided a foundation for understanding speech perception, but the introduction of spectrotemporal modulation may enhance predictive accuracy. The presentation first introduces speech intelligibility prediction models that utilize frequency and temporal information analysis mechanisms, such as auditory and modulation filterbanks. Next, prediction models that employ spectrotemporal modulation and machine learning-based models using these as features are discussed. Finally, the presentation highlights trends in the Clarity Project, which leverages large-scale listening experiment data for hearing-impaired listeners, along with recent state-of-the-art approaches based on speech foundation models. These studies are expected to contribute not only to the development of engineering theoretical models but also to practical applications in hearing assistive technology and auditory diagnostics.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Katsuhiko Yamamoto (2025) studied this question.

synapsesocial.com/papers/6a0567d2a550a87e60a1ffd6https://doi.org/10.1121/10.0040949
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Evaluating a 3-factor listener model for prediction of speech intelligibility to hearing-impaired listeners2024
  2. 2Predicting spectro-temporal modulation detection thresholds of individual listeners with hearing loss: Physiological modeling based on dynamic pitch cues2025
  3. 3The Predictive Role of Specific Audiometric Frequencies in Speech-in-Noise Perception: A Machine Learning Approach2026
  4. 4Speech intelligibility prediction based on a physiological model of the human ear and a hierarchical spiking neural network2024 · 3 citations
  5. 5Adaptation to noise in spectrotemporal modulation detection and its relationship to word recognition2025