Abstract:In recent decades, there has been a great deal of research into the problem of bilingual speech recognition - to develop a recognizer that can handle inter- and intra-sentential language switching between two languages. This paper presents our recent work on the development of a grammar-constrained, Mandarin-English bilingual Speech Recognition System (MESRS) for real world music retrieval. Two of the main difficult issues in handling the bilingual speech recognition systems for real world applications are tackled in this paper. One is to balance the performance and the complexity of the bilingual speech recognition system; the other is to effectively deal with the matrix language accents in embedded language**. In order to process the intra-sentential language switching and reduce the amount of data required to robustly estimate statistical models, a compact single set of bilingual acoustic models derived by phone set merging and clustering is developed instead of using two separate monolingual models for each language. In our study, a novel Two-pass phone clustering method based on Confusion Matrix (TCM) is presented and compared with the log-likelihood measure method. Experiments testify that TCM can achieve better performance. Since potential system users' native language is Mandarin which is regarded as a matrix language in our application, their pronunciations of English as the embedded language usually contain Mandarin accents. In order to deal with the matrix language accents in embedded language, different non-native adaptation approaches are investigated. Experiments show that model retraining method outperforms the other common adaptation methods such as Maximum A Posteriori (MAP). With the effective incorporation of approaches on phone clustering and non-native adaptation, the Phrase Error Rate (PER) of MESRS for English utterances was reduced by 24.47% relatively compared to the baseline monolingual English system while the PER on Mandarin utterances was comparable to that of the baseline monolingual Mandarin system. The performance for bilingual utterances achieved 22.37% relative PER reduction.

On Continuous Speech Recognition of Indian English

[The Neural Encoding of Continuous Speech - Recent Advances in EEG and MEG Studies].

A study on native American English speech recognition by Indian listeners with varying word familiarity level

Deep Speech Based End-to-End Automated Speech Recognition (ASR) for Indian-English Accents

Speech Recognition With No Speech Or With Noisy Speech Beyond English

Improving EEG based Continuous Speech Recognition

Continuous Silent Speech Recognition using EEG

Cross-Corpora Language Recognition: A Preliminary Investigation with Indian Languages

Svarah: Evaluating English ASR Systems on Indian Accents

Study of Indian English Pronunciation Variabilities relative to Received Pronunciation

English Speech Recognition System Model Based on Computer-Aided Function and Neural Network Algorithm

Towards Robust Speech Recognition in the Telephony Network Environment - Cellular and Landline Conditions

Towards Building ASR Systems for the Next Billion Users

An Investigation of Indian Native Language Phonemic Influences on L2 English Pronunciations

Low-resource speech recognition and dialect identification of Irish in a multi-task framework

Development of A Mandarin-English Bilingual Speech Recognition System for Real World Music Retrieval

Toward On-Line Learning of Chinese Continuous Speech Recognition System.

A comparative study on selecting acoustic modeling units in deep neural networks based large vocabulary Chinese speech recognition

Improving Speech Recognition for African American English With Audio Classification

Innovative Application of Sensor Combined with Speech Recognition Technology in College English Education in the Context of Artificial Intelligence

Deep Neural Network-based Mixed Speech Recognition Technology for Chinese and English