Candidate Expansion Algorithm Based on Weighted Syllable Confusion Matrix for Mandarin LVCSR

Chang Fengxiang,Li Baoxiang,Liu Gang,Guo Jun
DOI: https://doi.org/10.1109/cc.2013.6571293
2013-01-01
China Communications
Abstract:The inclusion of more potentially correct words in the candidate sets is important to improve the accuracy of Large Vocabulary Continuous Speech Recognition (LVCSR). A candidate expansion algorithm based on the Weighted Syllable Confusion Matrix (WSCM) is proposed. First, WSCM is derived from a confusion network. Then, the recognised candidates in the confusion network is used to conjecture the most likely correct words based on WSCM, after which, the conjectured words are combined with the recognised candidates to produce an expanded candidate set. Finally, a combined model having mutual information and a trigram language model is used to rerank the candidates. The experiments on Mandarin film data show that an improvement of 9.57% in the character correction rate is obtained over the initial recognition performance on those light erroneous utterances.
What problem does this paper attempt to address?