Abstract:End-to-end attention-based modeling is increasingly popular for tackling sequence-to-sequence mapping tasks. Traditional attention mechanisms utilize prior input information to derive attention, which then conditions the output. However, we believe that knowledge of posterior output information may convey some advantage when modeling attention. A recent technique proposed for machine translation called the posterior attention model (PAM) demonstrates that posterior output information can be used in that way for machine translation. This paper explores the use of posterior information for attention modeling in an automatic speech recognition (ASR) task. We demonstrate that direct application of PAM to ASR is unsatisfactory, due to two deficiencies; Firstly, PAM adopts attention based weighted single-frame output prediction by assuming a single focused attention variable, whereas wider contextual information from acoustic frames is important for output prediction in ASR. Secondly, in addition to the well-known exposure bias problem, PAM introduces additional mismatches in attention training and inference calculations. We present extensive experiments combining a number of alternative approaches to solving these problems, leading to a high performance technique which we call extended PAM (EPAM). To counter the first deficiency, EPAM modifies the encoder to introduce additional context information for output prediction. The second deficiency is overcome in EPAM through a two part solution of a mismatch penalty term and an alternate learning strategy. The former applies a divergence-based loss to correct the mismatch bias distribution, while the latter employs a novel update strategy which relies on introducing iterative inference steps alongside each training step. In experiments with both WSJ-80hrs and Switchboard-300hrs datasets we found significant performance gains. For example, the full EPAM system model achieved a word error rate (WER) of 10.6% on the WSJ eval92 test set, compared to 11.6% for traditional prior-attention modeling. Meanwhile, on the Switchboard eval2000 test set, we achieved 16.3% WER compared to the traditional method WER of 17.3%.

Gaussian Prediction Based Attention for Online End-to-End Speech Recognition.

Segment Boundary Detection Directed Attention for Online End-to-end Speech Recognition

An Online Attention-based Model for Speech Recognition

Efficient Decoding Self-Attention for End-to-end Speech Synthesis

Online Hybrid CTC/Attention End-to-End Automatic Speech Recognition Architecture

Attention-based Transducer for Online Speech Recognition

Learning Adaptive Downsampling Encoding for Online End-to-End Speech Recognition

Deep Latent Variable Predictive Modeling with Online Bayesian Soft Attention Mechanism

CGA-MGAN: Metric GAN Based on Convolution-Augmented Gated Attention for Speech Enhancement

Monotonic Gaussian regularization of attention for robust automatic speech recognition

Attention does not guarantee best performance in speech enhancement

An Attention-Based Neural Network Approach For Single Channel Speech Enhancement

Improving Attention-Based End-to-End Speech Recognition by Monotonic Alignment Attention Matrix Reconstruction.

Attention-based Recurrent Generator with Gaussian Tolerance for Statistical Parametric Speech Synthesis

Parallel Gated Neural Network With Attention Mechanism For Speech Enhancement

Improving Speech Recognition Error Prediction for Modern and Off-the-shelf Speech Recognizers

Effective Exploitation of Posterior Information for Attention-Based Speech Recognition.

Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech Recognition

Transformer-based end-to-end speech recognition with residual Gaussian-based self-attention

Online Speaker Adaptation for LVCSR Based on Attention Mechanism

Improved Multi-Stage Training of Online Attention-based Encoder-Decoder Models