Binaural Selective Attention Model for Target Speaker Extraction

Hanyu Meng,Qiquan Zhang,Xiangyu Zhang,Vidhyasaharan Sethu,Eliathamby Ambikairajah
2024-06-18
Abstract:The remarkable ability of humans to selectively focus on a target speaker in cocktail party scenarios is facilitated by binaural audio processing. In this paper, we present a binaural time-domain Target Speaker Extraction model based on the Filter-and-Sum Network (FaSNet). Inspired by human selective hearing, our proposed model introduces target speaker embedding into separators using a multi-head attention-based selective attention block. We also compared two binaural interaction approaches -- the cosine similarity of time-domain signals and inter-channel correlation in learned spectral representations. Our experimental results show that our proposed model outperforms monaural configurations and state-of-the-art multi-channel target speaker extraction models, achieving best-in-class performance with 18.52 dB SI-SDR, 19.12 dB SDR, and 3.05 PESQ scores under anechoic two-speaker test configurations.
Audio and Speech Processing,Sound,Signal Processing
What problem does this paper attempt to address?
The goal of this paper is to address the problem of simulating the human ability to selectively attend to a specific speaker (target speaker) in complex acoustic environments (such as cocktail party scenarios) through binaural hearing. Specifically, the research focuses on developing a model based on Binaural Time-Domain Target Speaker Extraction (TSE), which is built upon the Filter-and-Sum Network (FaSNet) and incorporates a selective attention module with a multi-head attention mechanism to adapt to speaker embeddings. Additionally, the paper explores two binaural interaction methods: cosine similarity of time-domain signals and inter-channel correlation in learned spectral representations, to further enhance model performance. The main contributions of the research include: 1. Proposing a binaural selective attention model that uses FaSNet as a foundation and simulates binaural input through Head-Related Transfer Functions (HRTFs). 2. Designing a selective attention block to adapt speaker embeddings into the separator. 3. Exploring two binaural interaction methods: Cosine Similarity (CSim) of time-domain signals and Inter-Channel Attention Correlation (IAC) in learned spectral representations, thereby proposing two binaural target speaker extraction models (Bi-CSim-TSE and Bi-IAC-TSE). 4. Experimental results show that the proposed models outperform the mono configuration and existing multi-channel target speaker extraction models in a two-speaker test configuration without echo, achieving the best level of performance metrics (18.52 dB SI-SDR, 19.12 dB SDR, and 3.05 PESQ score).