Efficient text feature extraction by integrating the average linkage and K-medoids clustering

Dasong Sun
DOI: https://doi.org/10.1142/S0217984921501517
2021-02-18
Modern Physics Letters B
Abstract:By clustering feature words, we can not only simplify the dimension of feature subsets, but also eliminate the redundancy of the feature. However, for a feature set with very large dimensions, the traditional K -medoids algorithm is difficult to accurately estimate the value of k . Moreover, the clustering results of the average linkage (AL) algorithm cannot be divided again, and the AL algorithm cannot be directly used for text classification. In order to overcome the limitations of AL and K -medoids, in this paper, we combine the two algorithms together so as to be mutually complementary to each other. In particular, in order to meet the purpose of text classification, we improve the AL algorithm and propose the R2 testing statistics to obtain the approximate number of clusters. Finally, the central feature words are preserved, and the other feature words are deleted. The experimental results show that the new algorithm largely eliminates the redundancy of the feature. Compared with the traditional TF-IDF algorithms, the performance of the text classification of the new algorithm is improved.
physics, condensed matter, applied, mathematical
What problem does this paper attempt to address?