Abstract:Although great progress has been made in automatic speech recognition (ASR), significant performance degradation still exists in noisy environments. In this paper, a novel factor-aware training framework, named neural network-based multifactor aware joint training, is proposed to improve the recognition accuracy for noise robust speech recognition. This approach is a structured model which integrates several different functional modules into one computational deep model. We explore and extract speaker, phone, and environment factor representations using deep neural networks (DNNs), which are integrated into the main ASR DNN to improve classification accuracy. In addition, the hidden activations in the main ASR DNN are used to improve factor extraction, which in turn helps theASRDNN. All the model parameters, including those in the ASR DNN and factor extraction DNNs, are jointly optimized under the multitask learning framework. Unlike prior traditional techniques for the factor-aware training, our approach requires no explicit separate stages for factor extraction and adaptation. Moreover, the proposed neural network-based multifactor aware joint training can be easily combined with the conventional factor-aware training which uses the explicit factors, such as i-vector, noise energy, and T 60 value to obtain additional improvement. The proposed method is evaluated on two main noise robust tasks: the AMI single distant microphone task in which reverberation is the main concern, and the Aurora4 task in which multiple noise types exist. Experiments on both tasks show that the proposed model can significantly reduce word error rate (WER). The best configuration achieved more than 15% relative reduction in WER over the baselines on these two tasks.

Uyghur speech recognition based on deep neural network

Mongolian acoustic modeling based on deep neural network

Deep Neural Network based Uyghur Large Vocabulary Continuous Speech Recognition

Mongolian Speech Recognition Based on Deep Neural Networks

Speech Recognition Based on Deep Neural Networks on Tibetan Corpus

Performance Optimization of Speech Recognition System with Deep Neural Network Model

A Parameter Transfer Method for HMM-DNN Heterogeneous Model with the Scarce Mongolian Data Set

Research on Improving Phoneme Recognition Rate Based on Subspace Gaussian Mixture Model and Deep Neural Network Combination

A Chinese Acoustic Model Based On Convolutional Neural Network

Speech Recognition Based on Deep Tensor Neural Network and Multifactor Feature.

Crosslingual Acoustic Modeling in Uyghur Speech Recognition

Acceleration Strategies for Speech Recognition Based on Deep Neural Networks

Robust Speech Recognition With Speech Enhanced Deep Neural Networks

State-Clustering Based Multiple Deep Neural Networks Modeling Approach for Speech Recognition

Research on Uyghur Speech Acoustic Recognition System Based on DDBHMM

Deep Neural Network-based Mixed Speech Recognition Technology for Chinese and English

Automatic Speech Recognition Method Based on Deep Learning Approaches for Uzbek Language

Phonotactic language recognition based on DNN-HMM acoustic model

DLD: An Optimized Chinese Speech Recognition Model Based on Deep Learning

Neural Network Based Multi-Factor Aware Joint Training for Robust Speech Recognition.

Speech Recognition Algorithm Based on Neural Network and Hidden Markov Model