Abstract:Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an open issue. Two main challenges are the complex acoustic environment and the real-time processing requirement. To address these challenges, we propose a temporal-spatial neural filter, which directly estimates the target speech waveform from multi-speaker mixture in reverberant environments, assisted with directional information of the speaker(s). Firstly, against variations brought by complex environment, the key idea is to increase the acoustic representation completeness through the jointly modeling of temporal, spectral and spatial discriminability between the target and interference source. Specifically, temporal, spectral, spatial along with the designed directional features are integrated to create a joint acoustic representation. Secondly, to reduce the latency, we design a fully-convolutional autoencoder framework, which is purely end-to-end and single-pass. All the feature computation is implemented by the network layers and operations to speed up the separation procedure. Evaluation is conducted on simulated reverberant dataset WSJ0-2mix and WSJ0-3mix under speaker-independent scenario. Experimental results demonstrate that the proposed method outperforms state-of-the-art deep learning based multi-channel approaches with fewer parameters and faster processing speed. Furthermore, the proposed temporal-spatial neural filter can handle mixtures with varying and unknown number of speakers and exhibits persistent performance even when existing a direction estimation error. Codes and models will be released soon.

3D Spatial Features for Multi-Channel Target Speech Separation

Boosting Spatial Information for Deep Learning Based Multichannel Speaker-Independent Speech Separation in Reverberant Environments.

Temporal-Spatial Neural Filter: Direction Informed End-to-End Multi-channel Target Speech Separation

Locate and Beamform: Two-dimensional Locating All-neural Beamformer for Multi-channel Speech Separation

Gated Recurrent Fusion of Spatial and Spectral Features for Multi-Channel Speech Separation with Deep Embedding Representations.

RIR-SF: Room Impulse Response Based Spatial Feature for Target Speech Recognition in Multi-Channel Multi-Speaker Scenarios

A Speaker-Dependent Approach to Separation of Far-Field Multi-Talker Microphone Array Speech for Front-End Processing in the CHiME-5 Challenge

Localization Based Stereo Speech Separation Using Deep Networks.

Robust Spatial Filtering Network for Separating Speech in the Direction of Interest

Multi-channel Speech Separation Using Spatially Selective Deep Non-linear Filters

A Study of Multichannel Spatiotemporal Features and Knowledge Distillation on Robust Target Speaker Extraction

Separating Voices from Multiple Sound Sources Using 2D Microphone Array

Adaptive Beamforming Based on Interference-Plus-Noise Covariance Matrix Reconstruction for Speech Separation

Speaker and Direction Inferred Dual-channel Speech Separation

A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition.

A Multi-channel Speech Separation System for Unknown Number of Multiple Speakers

3-D Feature and Acoustic Modeling for Far-Field Speech Recognition

Dual-Channel Speech Separation by Sub-Segmental Directional Statistics

Deep Ad-hoc Beamforming Based on Speaker Extraction for Target-Dependent Speech Separation

A Channel-Wise Multichannel Speech Separation Network for Ad-hoc Microphones Environments

Challenges and Insights: Exploring 3D Spatial Features and Complex Networks on the MISP Dataset