A new word detection method for chinese based on local context information

ZENG Hua-lin,ZHOU Chang-le,ZHENG Xu-ling
2010-01-01
Abstract:Finding out out-of-vocabulary words is an urgent and difficult task in Chinese words segmentation. To avoid the defect causing by offline training in the traditional method, the paper proposes an Improved prediction by partical match (PPM) segmenting algorithm for Chinese words based on extracting local context information, which adds the context information of the testing text into the local PPM statistical model so as to guide the detection of new words. The algorithm focuses on the process of online segmentation and new word detection which achieves a good effect in the close or opening test, and outperforms some well-known Chinese segmentation system to a certain extent. new word detection, improved PPM model, context information, Chinese words segmentation.
What problem does this paper attempt to address?