An Improved TF-IDF Weights Function Based on Information Theory

Na Wang,Pengyuan Wang,Baowei Zhang
DOI: https://doi.org/10.1109/cctae.2010.5544382
2010-01-01
Abstract:Vector Space Model (VSM) is a typical method to describe the text feature in text classification at present. It adopts TF-IDF weights to compute the term weighting in each dimension of the text feature. However, it only considers the relationship between the term and the whole text but neglects the relationship between different terms. Aiming at this problem an improved TF-IDF weights function is proposed which uses the distribution information among classes and inside a class. The experience shows that the improved method is feasible and effective. In addition, it greatly improves the accuracy of text category.
What problem does this paper attempt to address?