Abstract:In recent decades, there has been a great deal of research into the problem of bilingual speech recognition - to develop a recognizer that can handle inter- and intra-sentential language switching between two languages. This paper presents our recent work on the development of a grammar-constrained, Mandarin-English bilingual Speech Recognition System (MESRS) for real world music retrieval. Two of the main difficult issues in handling the bilingual speech recognition systems for real world applications are tackled in this paper. One is to balance the performance and the complexity of the bilingual speech recognition system; the other is to effectively deal with the matrix language accents in embedded language**. In order to process the intra-sentential language switching and reduce the amount of data required to robustly estimate statistical models, a compact single set of bilingual acoustic models derived by phone set merging and clustering is developed instead of using two separate monolingual models for each language. In our study, a novel Two-pass phone clustering method based on Confusion Matrix (TCM) is presented and compared with the log-likelihood measure method. Experiments testify that TCM can achieve better performance. Since potential system users' native language is Mandarin which is regarded as a matrix language in our application, their pronunciations of English as the embedded language usually contain Mandarin accents. In order to deal with the matrix language accents in embedded language, different non-native adaptation approaches are investigated. Experiments show that model retraining method outperforms the other common adaptation methods such as Maximum A Posteriori (MAP). With the effective incorporation of approaches on phone clustering and non-native adaptation, the Phrase Error Rate (PER) of MESRS for English utterances was reduced by 24.47% relatively compared to the baseline monolingual English system while the PER on Mandarin utterances was comparable to that of the baseline monolingual Mandarin system. The performance for bilingual utterances achieved 22.37% relative PER reduction.

An undergraduate Mandarin speech database for speaker recognition research

Preliminary Study on Self-contained UBM Construction for Speaker Recognition.

Design and implementation of a speaker recognition system

Pvd: A New Pathological Voice Dataset For Intra-Speaker Recognition Research Interest

A Universal Phoneme-Set Based Language Independent Short Utterance Speaker Recognition

MASC: A Speech Corpus in Mandarin for Emotion Analysis and Affective Speaker Recognition

KeSpeech: an Open Source Speech Dataset of Mandarin and Its Eight Subdialects.

High Performance Digit Mandarin Speech Recognition

AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection

Construction and Evaluation of Mandarin Multimodal Emotional Speech Database

ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5

TRSD: A Time-Varying and Region-Changed Speech Database for Speaker Recognition

The Mandarin Chinese speech database: a corpus of 18,820 auditory neutral nonsense sentences

CDSD: Chinese Dysarthria Speech Database

OC16-CE80: A Chinese-English Mixlingual Database and A Speech Recognition Baseline

Wavoice: A mmWave-assisted Noise-resistant Speech Recognition SystemJust Accepted

Development of A Mandarin-English Bilingual Speech Recognition System for Real World Music Retrieval

Enhancing Voice Wake-Up for Dysarthria: Mandarin Dysarthria Speech Corpus Release and Customized System Design

AISHELL-4 - An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

Improving Short Utterance Speaker Recognition by Modeling Speech Unit Classes

THCHS-30 : A Free Chinese Speech Corpus