PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 2024IEEE/ACM Transactions on Audio Speech and Language Processing28 citationsOpen Access

Speech Enhancement for Cochlear Implant Recipients Using Deep Complex Convolution Transformer With Frequency Transformation

NMNursadul MamunJHJohn H. L. Hansen

Key Points

Key points are not available for this paper at this time.

Abstract

The presence of background noise or competing talkers is one of the main communication challenges for cochlear implant (CI) users in speech understanding in naturalistic spaces. These external factors distort the time-frequency (T-F) content including magnitude spectrum and phase of speech signals. While most existing speech enhancement (SE) solutions focus solely on enhancing the magnitude response, recent research highlights the importance of phase in perceptual speech quality. Motivated by multi-task machine learning, this study proposes a deep complex convolution transformer network (DCCTN) for complex spectral mapping, which simultaneously enhances the magnitude and phase responses of speech. The proposed network leverages a complex-valued U-Net structure with a transformer within the bottleneck layer to capture sufficient low-level detail of contextual information in the T-F domain. To capture the harmonic correlation in speech, DCCTN incorporates a frequency transformation block in the encoder structure of the U-Net architecture. The DCCTN learns a complex transformation matrix to accurately recover speech in the T-F domain from a noisy input spectrogram. Experimental results demonstrate that the proposed DCCTN outperforms existing model solutions such as the convolutional recurrent network (CRN), deep complex convolutional recurrent network (DCCRN), and gated convolutional recurrent network (GCRN) in terms of objective speech intelligibility and quality, both for seen and unseen noise conditions. To evaluate the effectiveness of the proposed SE solution, a formal listener evaluation involving four CI recipients was conducted. Results indicate a significant improvement in speech intelligibility performance for CI recipients in noisy environments. Additionally, DCCTN demonstrates the capability to suppress highly non-stationary noise without introducing musical artifacts commonly observed in conventional SE methods.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mamun et al. (2024) studied this question.

synapsesocial.com/papers/6a152dacd64fa333899f54a9https://doi.org/10.1109/taslp.2024.3366760
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Cochlear Implants: System Design, Integration, and Evaluation2008 · 891 citations
  2. 2CCi-CLOUD: A framework for community based remote cochlear implant user experiments based on the CCi-MOBILE research platform2022 · 1 citations
  3. 3Funnel Deep Complex U-Net for Phase-Aware Speech Enhancement2021 · 10 citations
  4. 4Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks2015 · 702 citations
  5. 5Learning Complex Spectral Mapping With Gated Convolutional Recurrent Networks for Monaural Speech Enhancement2019 · 360 citations