This study addresses the challenge of selective auditory attention in noisy environments by proposing an EEG-based target speaker extraction model, ASEAF, designed to mimic neural decoding through tailored spatio-temporal feature extraction and cross-modal fusion. The model achieves precise extraction of the target speaker's speech by simultaneously processing EEG and audio signals. ASEAF comprises four modules: an EEG encoder using CNN and self-attention for spatio-temporal features, an audio encoder with SincNet for frequency-aware processing, a dual-path LSTM speaker extractor for fused feature masking, and a CNN decoder for waveform reconstruction. This innovative integration advances neural-signal-based speech reconstruction by providing insights into cross-modal interactions. Experiments on the Cocktail Party dataset, KUL dataset and DTU dataset demonstrate that ASEAF outperforms state-of-the-art models across multiple metrics, with an average improvement of 11.5% in scale-invariant signal-to-distortion ratio (SI-SDRi). This work offers a more effective hearing aid solution for individuals with hearing impairments and advances the field of brain-computer interfaces.
Yang et al. (Fri,) studied this question.