Syllable-Based Acoustic Modeling with CTC for Multi-Scenarios Mandarin Speech Recognition

Yuanyuan Zhao,Linhao Dong,Shuang Xu,Bo Xu
DOI: https://doi.org/10.1109/ijcnn.2018.8489589
2018-01-01
Abstract:With the improvement of speech recognition, voice products are gradually applied to every scene of life. The existing approaches to handle various scenarios are often to build many different acoustic models using scenario-dependent data only, with each for a special scene. The obvious weakness of these approaches is that it seriously hampers the large-scale application and maintenance of voice products. To address this issue, acoustic modeling based on context-independent syllables optimized with CTC loss is presented for multiple scenarios of Mandarin speech recognition. On the one hand, context-independent modeling overcomes the shortcomings of context-dependent modeling overfitting a particular scene. Also, it sidesteps decision trees used in context-dependent modeling so that there is no need to consider the building of decision tree and whether to start training again in a real application. On the other hand, choosing longer-length syllable acoustic units can effectively preserve the co-articulation effect that context-dependent phone can model. Also, syllables in the Chinese language have its inherent advantages, as its number is fixed and it is trainable, effective generalization and better robustness. This paper also explores the differences between wideband and narrowband data caused by the front-end signal acquisition block, and proposes a unified training method based on the use of VGG in the bottom layer, and introduces layer normalization. The experimental results demonstrate that the proposed syllable-based CTC acoustic model for multiple scenarios can achieve more than 15% and 7% relatively improvement for mobile phone data and telephone data separately compare with scenarios-dependent modeling.
What problem does this paper attempt to address?