Abstract:Non-native speech causes automatic speech recognition systems to degrade in performance. Past strategies to address this challenge have considered model adaptation, accent classification with a model selection, alternate pronunciation lexicon, etc. In this study, we consider a recurrent neural network (RNN) with connectionist temporal classification (CTC) cost function trained on multi-accent English data including US (Native), Indian and Hispanic accents. We exploit dark knowledge from a model trained with the multi-accent data to train student models under the guidance of both a teacher model and CTC cost of target transcription. We show that transferring knowledge from a single RNN-CTC trained model toward a student model, yields better performance than the stand-alone teacher model. Since the outputs of different trained CTC models are not necessarily aligned, it is not possible to simply use an ensemble of CTC teacher models. To address this problem, we train accent specific models under the guidance of a single multi-accent teacher, which results in having multiple aligned and trained CTC models. Furthermore, we train a student model under the supervision of the accent-specific teachers, resulting in an even further complementary model, which achieves +20.1% relative Character Error Rate (CER) reduction compared to the baseline trained without any teacher. Having this effective multi-accent model, we can achieve further improvement for each accent by adapting the model to each accent. Using the accent specific model's outputs to regularize the adapting process (i.e., a knowledge distillation version of Kullback-Leibler (KL) divergence) results in even superior performance compared to the conventional approach using general teacher models.

Capitalization Normalization for Language Modeling with an Accurate and Efficient Hierarchical RNN Model

Position-Invariant Truecasing with a Word-and-Character Hierarchical Recurrent Neural Network

reproducing "ner and pos when nothing is capitalized"

An Efficient Architecture for Predicting the Case of Characters using Sequence Models

Fast and Accurate Capitalization and Punctuation for Automatic Speech Recognition Using Transformer and Chunk Merging

A Character-Level Approach to the Text Normalization Problem Based on a New Causal Encoder

Hierarchical LSTM with char-subword-word tree-structure representation for Chinese named entity recognition

BiCAPT: Bidirectional Computer-Assisted Pronunciation Training with Normalizing Flows

Hierarchical Recurrent Neural Network for Document Modeling.

Normalizing Text using Language Modelling based on Phonetics and String Similarity

Multi-Level Modeling Units for End-to-End Mandarin Speech Recognition

Recurrent Neural Network Language Model With Structured Word Embeddings For Speech Recognition

Text normalization using memory augmented neural networks

Label-Correction Capsule Network for Hierarchical Text Classification.

Hierarchical Convolutional Recurrent Neural Network for Chinese Text Classification

Text Injection for Capitalization and Turn-Taking Prediction in Speech Models

A Character-Level Method for Text Classification

Transformer-based Models of Text Normalization for Speech Applications

Advancing Multi-Accented LSTM-CTC Speech Recognition using a Domain Specific Student-Teacher Learning Paradigm

Improving Vietnamese Named Entity Recognition from Speech Using Word Capitalization and Punctuation Recovery Models

An End-to-end Chinese Text Normalization Model Based on Rule-guided Flat-Lattice Transformer.