Abstract:The feature selection is an important part in automatic text classification. In this paper, we use a Chinese semantic dictionary-Hownet to extract the concepts from the word as the feature set, because it can better reflect the meaning of the text. However, as the concept definition in the dictionary sometimes cannot express the word properly, we define the expression power for every sememe and every definition of the word in further process, and define the relation degree between the sememe and the definition. A threshold is set in the sememe tree, the sememe of the little information is filtered, and the words of weak definition are reserved in expression power. By this method, we construct a combined feature set that consists of both sememes and the Chinese words. The values of sememes are given according to their expression power and relation to the word. By comparing seven feature weighing methods in text classification, we propose a CHI-MCOR weighing method according to the weighing theories and classification precision. Experimental result shows that if the words are extracted properly, not only the feature dimension is smaller but also the classification precision is higher. Our method makes a good balance between the features which occur frequently in the corpus and those which only occur in one category, the difference of the classification precision among different categories is small.

Concept Acquisition from Corpora: Using an Automatic Clustering Method Based on Chinese Measure Words

A New Hypred Improved Method for Measuring Concept Semantic Similarity in WordNet.

A Novel Comprehensive Approach for Estimating Concept Semantic Similarity in WordNet

Automatic Multi-Way Domain Concept Hierarchy Construction from Customer Reviews

Concept chain based text clustering

Word frequency approximation for chinese using raw, MM-Segmented and manually segmented corpora

A Comparative Study on Chinese Word Clustering

A Comparative Study on Chinese Word Segmentation Using Statistical Models

Chinese Word Frequency Approximation Based on Multitype Corpora.

A Multi-Concept Semantic Representation System for Chinese Intent Recognition

A Comparative Study on Representing Units in Chinese Text Clustering

Two-Character Chinese Word Extraction Based on Hybrid of Internal and Contextual Measures

Discovering Chinese Concept-In-Corpus

Research on the Semantic Measurement in Co-word Analysis.

A New Feature Selection Method Based On Concept Extraction In Automatic Chinese Text Classification

A morphology-based Chinese word segmentation method

Measuring Word Polysemousness And Sense Granularity At A Language Level

A Clustering Algorithm for Short Documents Based On Concept Similarity

A Fast and Effective Method for Clustering Large-Scale Chinese Question Dataset

Experimental Study On Representing Units In Chinese Text Categorization

Chinese WSD Based on Context Calculation Model