Abstract:A set of IR experiments was carried out to study the impact of Chinese word segmentation and its effect on information retrieval (IR) at the Division of Information Studies, Nanyang Technological University, Singapore. A total of four automatic character-based segmentation approaches and a manual word segmentation approach was first carried out to obtain the word segments for indexing and to evaluate the segmentation accuracy of these automatic approaches. The IR experiments study both the influence of different document segmentation approaches on IR effectiveness and the methods used for query segmentation. Traditional data recall and precision measures were used to gauge IR effectiveness. A number of queries were selected and subjected to further detailed analysis to further explore the influence of word segmentation on IR.The findings reveal that the segmentation approach has an effect on IR effectiveness. Better IR results are obtained by using the same method for query and document processing as this increase the probability of the query-document match. The recognition of a higher number of 2-character words generally contributes to the improvement of IR effectiveness. However, manual segmentation does not always work better than character-based segmentation as a result of the existence of longer words with more than two characters. No evidence is found that ambiguous words resulting from the segmentation process significantly affect IR.

Segmentation of Chinese Discourse in Content-Based Information Retrieval.

On Generalized-Topic-Based Chinese Discourse Structure.

Chinese Discourse Segmentation Using Bilingual Discourse Commonality

Chinese Word Segmentation Method Based on Dictionary and Frequency of the Words

New Word Identification in Social Network Text Based on Time Series Information

The Locations of Word Segmentation in Chinese Reading: Research Based on the Eye-Movement-Contingent Display Technique

Intelligent Segmentation Framework and Data Hierarchy of Chinese Language and Literature Based on Semantic Recognition

A Statistical Approach For Resolving Problematical Word Boundaries In Chinese Lexicography

TopWORDS-Seg: Simultaneous Text Segmentation and Word Discovery for Open-Domain Chinese Texts via Bayesian Inference

Chinese word segmentation and its effect on information retrieval

Topic Detection Technology for Chinese Text Based on Statistics and Semantic Information

A Survey of Chinese Word Segmentation and the Application in Information Retrieval

Unsupervised segmentation of chinese corpus using accessor variety

A Pragmatic Approach for Classical Chinese Word Segmentation.

An automatic approach for efficient text segmentation

A Discriminative Latent Variable Chinese Segmenter with Hybrid Word/Character Information.

Toward Fast and Accurate Neural Discourse Segmentation

Knowledge-Based Approaches to the Segmentation of Oral History Interviews

Learning Recursive Segments for Discourse Parsing

Segmentation standard for Chinese natural language processing

Survey on Chinese Word Segmentation