Document Structure Analysis and Text Normalization for Chinese Putonghua and Cantonese Text-to-Speech Synthesis

Zhou Xinxin,Wu Zhiyong,Yuan Chun,Zhong Yuzhuo
DOI: https://doi.org/10.1109/IITA.2008.28
2008-01-01
Abstract:This paper describes our recent effort on document structure analysis (DSA) and text normalization (NORM) for Chinese Putonghua and Cantonese text-to-speech synthesis. A unified framework has been proposed, where DSA and NORM procedures are language-independent for the two-dialects of Chinese. For document structure analysis, regular expressions have been utilized to detect and identify the non-standard-words (NSWs) and punctuations related to document structure; a new document segmentation approach is then proposed by considering the information provided by NSWs and punctuations. For text normalization, a method which considers the contextual information is put forward to handle the ambiguity of the NSWs, symbols and punctuations.
What problem does this paper attempt to address?