Abstract:Non-native speech causes automatic speech recognition systems to degrade in performance. Past strategies to address this challenge have considered model adaptation, accent classification with a model selection, alternate pronunciation lexicon, etc. In this study, we consider a recurrent neural network (RNN) with connectionist temporal classification (CTC) cost function trained on multi-accent English data including US (Native), Indian and Hispanic accents. We exploit dark knowledge from a model trained with the multi-accent data to train student models under the guidance of both a teacher model and CTC cost of target transcription. We show that transferring knowledge from a single RNN-CTC trained model toward a student model, yields better performance than the stand-alone teacher model. Since the outputs of different trained CTC models are not necessarily aligned, it is not possible to simply use an ensemble of CTC teacher models. To address this problem, we train accent specific models under the guidance of a single multi-accent teacher, which results in having multiple aligned and trained CTC models. Furthermore, we train a student model under the supervision of the accent-specific teachers, resulting in an even further complementary model, which achieves +20.1% relative Character Error Rate (CER) reduction compared to the baseline trained without any teacher. Having this effective multi-accent model, we can achieve further improvement for each accent by adapting the model to each accent. Using the accent specific model's outputs to regularize the adapting process (i.e., a knowledge distillation version of Kullback-Leibler (KL) divergence) results in even superior performance compared to the conventional approach using general teacher models.

Unsupervised Regularization-Based Adaptive Training for Speech Recognition

Adaptive Speaker Normalization for CTC-Based Speech Recognition

Speaker Adaptation for End-to-End CTC Models.

CTC Regularized Model Adaptation for Improving LSTM RNN Based Multi-Accent Mandarin Speech Recognition

AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition

CR-CTC: Consistency regularization on CTC for improved speech recognition

Attention-Based Gated Scaling Adaptive Acoustic Model for CTC-Based Speech Recognition.

Improving CTC-AED model with integrated-CTC and auxiliary loss regularization

Speaker Normalization Training and Adaptation for Speech Recognition

Nonlinear Regularization Decoding Method for Speech Recognition

Speaker-Invariant Training Via Adversarial Learning.

Advancing Multi-Accented LSTM-CTC Speech Recognition using a Domain Specific Student-Teacher Learning Paradigm

Learning from Flawed Data: Weakly Supervised Automatic Speech Recognition

Dual-model self-regularization and fusion for domain adaptation of robust speaker verification

Improved Regularization Techniques for End-to-End Speech Recognition

Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC

Residual Convolutional CTC Networks for Automatic Speech Recognition.

KL-Divergence Regularized Deep Neural Network Adaptation for Low-Resource Speaker-Dependent Speech Enhancement

Unimodal Aggregation for CTC-based Speech Recognition

Acoustic Modeling With Dfsmn-Ctc And Joint Ctc-Ce Learning

Speaker Adaptation of Hybrid NN/HMM Model for Speech Recognition Based on Singular Value Decomposition