Key points are not available for this paper at this time.
Leveraging the predictive capabilities of neural networks to estimate intermediate parameters to derive a linear filter has proven to be an effective and robust approach to multi-channel speech enhancement. Here, we adopt the parameterized Variable Span Linear Filter (VSLF) framework, which provides a generalization of classical beamformer designs like the Multi-Channel Wiener Filter or Minimum Variance Distortionless Response. We propose an architecture that uses a deep neural network (DNN) to predict three intermediate quantities – both the clean-speech and overall-noise spatial covariance matrices, and an optimal speech-distortion tradeoff parameter – to compute the complex VSLF weights. Experiments show that our method improves speech intelligibility and quality over end-to-end baselines, while enabling explicit control of the tradeoff between speech distortion and noise reduction.
Oviste et al. (Tue,) studied this question.