Using Duration and Pitch for Mandarin Digit String Recognition.

Rui Zhao,Yusuke Kida,Xiang Yan,Pei Ding,Lei He
DOI: https://doi.org/10.1109/icassp.2010.5495128
2010-01-01
Abstract:Mandarin digit string recognition (MDSR) is a challenge because there exist many difficulties in acoustic discrimination for such a small vocabulary speech recognition task. In this paper, we propose to improve MDSR performance by using duration and pitch information. Speech rate dispersion is used to involve duration knowledge and is incorporated in the MDSR system by rescoring the N-best candidates in a two-pass framework. estimated with a robust pitch extraction method is also adopted to improve the acoustic discrimination among Mandarin digits. The experimental results show both duration and pitch significantly improve the performance, and the combination of them gives further improvement. Moreover, our methods are robust to background noise. In the evaluation, the sentence error rate is reduced by 50.43% on average over different SNR conditions.
What problem does this paper attempt to address?