Abstract:This article presents a pragmatic approach to Chinese word segmentation. It differs from most previous approaches mainly in three respects. First, while theoretical linguists have defined Chinese words using various linguistic criteria, Chinese words in this study are defined pragmatically as segmentation units whose definition depends on how they are used and processed in realistic computer applications. Second, we propose a pragmatic mathematical framework in which segmenting known words and detecting unknown words of different types (i.e., morphologically derived words, factoids, named entities, and other unlisted words) can be performed simultaneously in a unified way. These tasks are usually conducted separately in other systems. Finally, we do not assume the existence of a universal word segmentation standard that is application-independent. Instead, we argue for the necessity of multiple segmentation standards due to the pragmatic fact that different natural language processing applications might require different granularities of Chinese words. These pragmatic approaches have been implemented in an adaptive Chinese word segmenter, called MSRSeg, which will be described in detail. It consists of two components: (1) a generic segmenter that is based on the framework of linear mixture models and provides a unified approach to the five fundamental features of word-level Chinese language processing: lexicon word processing, morphological analysis, factoid detection, named entity recognition, and new word identification; and (2) a set of output adaptors for adapting the output of (1) to different application-specific standards. Evaluation on five test sets with different standards shows that the adaptive system achieves state-of-the-art performance on all the test sets.

A Pragmatic Approach for Classical Chinese Word Segmentation.

Chinese Word Segmentation and Named Entity Recognition: A Pragmatic Approach

Chinese Word Segmentation Method Based on Dictionary and Frequency of the Words

A realistic and robust model for Chinese word segmentation

Chinese Word Segmentation Without Using Lexicon and Hand-Crafted Training Data

Chinese word segmentation at Peking University

Survey on Chinese Word Segmentation

A Comparison Study of Candidate Generation for Chinese Word Segmentation

A New Error-driven Learning Approach for Chinese Word Segmentation

A Statistical Approach For Resolving Problematical Word Boundaries In Chinese Lexicography

Adversarial Multi-Criteria Learning for Chinese Word Segmentation

RethinkCWS: is Chinese Word Segmentation a Solved Task?

Feature Abstraction for Lightweight and Accurate Chinese Word Segmentation

Towards Unified Chinese Segmentation Algorithm

When Classical Chinese Meets Machine Learning: Explaining the Relative Performances of Word and Sentence Segmentation Tasks

Unsupervised Chinese Word Segmentation with BERT Oriented Probing and Transformation

Word Segmentation for Classical Chinese Buddhist Literature

A New Psychometric-inspired Evaluation Metric for Chinese Word Segmentation.

Unsupervised segmentation of chinese corpus using accessor variety

Chinese Word Segmentation with Character Abstraction.

A Compression-based Algorithm for Chinese Word Segmentation