Using Gini-Index for Feature Selection in Text Categorization

LIN Yong-min,ZHU Wei-dong
DOI: https://doi.org/10.2991/icibet-14.2014.22
2007-01-01
Journal of Computer Applications
Abstract:With the rapid development of World Wide Web, text categorization has played an important role in organizing and processing large amount of text data.The first and major problem of text categorization is how to select the best subset from the original high feature space in order to reduce the high dimensionality of the original feature space and improve the classification performance.We aim to use improved Gini-index for text feature selection, constructing the measure function based on Gini-Index.We compare it to other four feature selection measures using two kinds of classifiers on two different document corpus.The result of experiments shows that its performance is comparable with other text feature selection approaches.However, it is perfect in the time complexity of algorithm.
What problem does this paper attempt to address?