Recently, speech extraction using a microphone array mounted on an unmanned aerial vehicle (UAV) has been attracting attention in situations such as rescue operations. However, the observation signal contains high levels of ego-noise, which is the noise generated by the UAV rotors, and the acoustic environments become very harsh. To achieve effective speech extraction in such environments, many methods utilizing frequency and spatial transfer characteristics of pre-recorded ego-noise have been proposed. One of them is a method that employs spatially regularized independent low-rank matrix analysis (SR-ILRMA) and rank-constrained spatial covariance matrix estimation (RCSCME). SR-ILRMA introduces spatial prior information into a blind source separation method, and RCSCME utilizes the outputs of SR-ILRMA to extract the target speech signal, resulting in a superior performance compared to other UAV speech extraction methods. In this paper, we introduce a noise prior consisting of pre-recorded ego-noise into RCSCME. For further performance improvement, we also extend the generative model of RCSCME, which is the multivariate complex Gaussian distribution in the conventional UAV speech extraction method, to the multivariate complex generalized Gaussian distribution. A simulated experiment confirms the effectiveness of the proposed approach in harsh environments, e.g., the input signal-to-noise ratio is −30 dB.
Nishikori et al. (Wed,) studied this question.