Abstract:Recent state-of-the-art (SOTA) effective neural network methods and fine-tuning methods based on pre-trained models (PTM) have been used in Chinese word segmentation (CWS), and they achieve great results. However, previous works focus on training the models with the fixed corpus at every iteration. The intermediate generated information is also valuable. Besides, the robustness of the previous neural methods is limited by the large-scale annotated data. There are a few noises in the annotated corpus. Limited efforts have been made by previous studies to deal with such problems. In this work, we propose a self-supervised CWS approach with a straightforward and effective architecture. First, we train a word segmentation model and use it to generate the segmentation results. Then, we use a revised masked language model (MLM) to evaluate the quality of the segmentation results based on the predictions of the MLM. Finally, we leverage the evaluations to aid the training of the segmenter by improved minimum risk training. Experimental results show that our approach outperforms previous methods on 9 different CWS datasets with single criterion training and multiple criteria training and achieves better robustness.(1)

Improving Chinese Word Segmentation Using Partially Annotated Sentences

Semi-Supervised Learning for Semantic Segmentation of Emphysema With Partial Annotations

Segment, Mask, and Predict: Augmenting Chinese Word Segmentation with Self-Supervision

Unsupervised Learning helps Supervised Neural Word Segmentation

Joint Chinese Word Segmentation and POS Tagging on Heterogeneous Annotated Corpora with Multiple Task Learning.

A Unified Model for Joint Chinese Word Segmentation and POS Tagging with Heterogeneous Annotation Corpora.

Enhancing Chinese Word Segmentation Using Unlabeled Data

A Sentence Segmentation Method for Ancient Chinese Texts Based on NNLM.

Neural Networks Incorporating Unlabeled and Partially-labeled Data for Cross-domain Chinese Word Segmentation

Deep Learning for Chinese Word Segmentation and POS Tagging.

Improving Chinese Word Segmentation on Micro-blog Using Rich Punctuations.

Chinese Word Segmentation Without Using Lexicon and Hand-Crafted Training Data

Parsing-based Chinese word segmentation integrating morphological and syntactic information

Exploring Representations from Unlabeled Data with Co-training for Chinese Word Segmentation.

Improving Cross-Domain Chinese Word Segmentation with Word Embeddings

Ancient Chinese Word Segmentation and Part-of-Speech Tagging Using Distant Supervision

Reducing Approximation and Estimation Errors for Chinese Lexical Processing with Heterogeneous Annotations

Neural Word Segmentation Learning for Chinese

Neural Chinese Word Segmentation with Lexicon and Unlabeled Data via Posterior Regularization

A Chinese Word Segmentation for Statistical Machine Translation

A Local Generative Model For Chinese Word Segmentation