Abstract:Learning an effective speaker representation is crucial for achieving reliable performance in speaker verification tasks. Speech signals are high-dimensional, long, and variable-length sequences containing diverse information at each time-frequency (TF) location. The standard convolutional layer that operates on neighboring local regions often fails to capture the complex TF global information. Our motivation is to alleviate these challenges by increasing the modeling capacity, emphasizing significant information, and suppressing possible redundancies. We aim to design a more robust and efficient speaker recognition system by incorporating the benefits of attention mechanisms and Discrete Cosine Transform (DCT) based signal processing techniques, to effectively represent the global information in speech signals. To achieve this, we propose a general global time-frequency context modeling block for speaker modeling. First, an attention-based context model is introduced to capture the long-range and non-local relationship across different time-frequency locations. Second, a 2D-DCT based context model is proposed to improve model efficiency and examine the benefits of signal modeling. A multi-DCT attention mechanism is presented to improve modeling power with alternate DCT base forms. Finally, the global context information is used to recalibrate salient time-frequency locations by computing the similarity between the global context and local features. This effectively improves the speaker verification performance compared to the standard ResNet model and Squeeze & Excitation block by a large margin. Our experimental results show that the proposed global context modeling method can efficiently improve the learned speaker representations by achieving channel-wise and time-frequency feature recalibration.

Improving ECAPA-TDNN Performance with Coordinate Attention

Two-dimensional discrete feature based spatial attention CapsNet For sEMG signal recognition

A Channel-Wise Spatial-Temporal Aggregation Network for Action Recognition

PCF: ECAPA-TDNN with Progressive Channel Fusion for Speaker Verification

ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks

Improving CTC-AED model with integrated-CTC and auxiliary loss regularization

Attention-Based Deep Spiking Neural Networks for Temporal Credit Assignment Problems.

Multi-Frequency Information Enhanced Channel Attention Module for Speaker Representation Learning

ELA: Efficient Local Attention for Deep Convolutional Neural Networks

CASE-Net: Integrating local and non-local attention operations for speech enhancement

Improving Acoustic Echo Cancellation by Exploring Speech and Echo Affinity with Multi-Head Attention.

A Symmetric Efficient Spatial and Channel Attention (ESCA) Module Based on Convolutional Neural Networks

Attention and DCT based Global Context Modeling for Text-independent Speaker Recognition

Improving Visual Speech Enhancement Network by Learning Audio-visual Affinity with Multi-head Attention

Target Speaker Extraction Using Attention-Enhanced Temporal Convolutional Network

NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification

Efficient Attention Network: Accelerate Attention by Searching Where to Plug

Neighborhood Attention Transformer with Progressive Channel Fusion for Speaker Verification

Efficient Multi-Scale Attention Module with Cross-Spatial Learning

Combining a parallel 2D CNN with a self-attention Dilated Residual Network for CTC-based discrete speech emotion recognition