Los puntos clave no están disponibles para este artículo en este momento.
Studies indicate that humans have the exceptional capability to focus on speech even in environments with complex background noises. To create a computational equivalent of this ability, recent studies have focused on developing speech extraction algorithms. These algorithms aim to extract human speech signals from a mixture of overlapping background sounds. In this paper, efforts are made to utilize the novel encoder-decoder neural network architecture to accomplish this task for real-time and streaming applications. The training of the model involved utilizing synthetic 6-second single-channel composite audio mixtures generated from three distinct datasets using the Scaper toolkit. Our results showcase an impressive SI-SNRi of 14.209 and a latency of 874 ms, obtained on a consumer-grade Intel Quad-Core i5 processor with default multithreading. The obtained outcomes provide compelling evidence for the feasibility and efficacy of the proposed approach, establishing a strong proof-of-concept for the real-world application of the model. Streaming human speech extraction could revolutionize acoustic applications for headphones, hearing aids, and telephony.
Nettam et al. (Thu,) studied this question.