Real-time speech extraction (SE) is an important task and has diverse applications, including pre-processing of dialogue systems. For wide applicability, it is preferable to reduce the latency of a real-time SE system as much as possible. We previously proposed real-time multichannel SE framework based on rank-constrained spatial covariance matrix estimation (RCSCME). In this framework, the observed signal is transformed into time-frequency domain using short-time Fourier transform (STFT), the target speech signal is extracted using RCSCME in a frame-by-frame manner, and the time-domain extracted speech signal is output by performing inverse STFT (ISTFT). As a result, the algorithmic latency becomes the length of the window function in STFT. However, to achieve sufficient SE performance, a short window length is inappropriate. In this research, we introduce well-known asymmetric window function technique, in which the length of the analysis window in STFT does not match that of the synthesis window in ISTFT, into our real-time SE framework. A long analysis window contributes to the high performance, while a short synthesis window reduces the algorithmic latency. In experiments, we demonstrate that the proposed method can significantly reduce the total latency compared with the conventional method while preserving the speech extraction performance.
Ishikawa et al. (Wed,) studied this question.