Abstract:Text classification, as an important research area of text mining, can quickly and effectively extract valuable information to address the challenges of organizing and managing large-scale text data in the era of big data. Currently, the related research on text classification tends to focus on the application in fields such as information filtering, information retrieval, public opinion monitoring, and library and information, with few studies applying text classification methods to the field of tourist attractions. In light of this, a corpus of tourist attraction description texts is constructed using web crawler technology in this paper. We propose a novel text representation method that combines Word2Vec word embeddings with TF-IDF-CRF-POS weighting, optimizing traditional TF-IDF by incorporating total relative term frequency, category discriminability, and part-of-speech information. Subsequently, the proposed algorithm respectively combines seven commonly used classifiers (DT, SVM, LR, NB, MLP, RF, and KNN), known for their good performance, to achieve multi-class text classification for six subcategories of national A-level tourist attractions. The effectiveness and superiority of this algorithm are validated by comparing the overall performance, specific category performance, and model stability against several commonly used text representation methods. The results demonstrate that the newly proposed algorithm achieves higher accuracy and F1-measure on this type of professional dataset, and even outperforms the high-performance BERT classification model currently favored by the industry. Acc, marco-F1, and mirco-F1 values are respectively 2.29%, 5.55%, and 2.90% higher. Moreover, the algorithm can identify rare categories in the imbalanced dataset and exhibit better stability across datasets of different sizes. Overall, the algorithm presented in this paper exhibits superior classification performance and robustness. In addition, the conclusions obtained by the predicted value and the true value are consistent, indicating that this algorithm is practical. The professional domain text dataset used in this paper poses higher challenges due to its complexity (uneven text length, relatively imbalanced categories), and a high degree of similarity between categories. However, this proposed algorithm can efficiently implement the classification of multiple subcategories of this type of text set, which is a beneficial exploration of the application research of complex Chinese text datasets in specific fields, and provides a useful reference for the vector expression and classification of text datasets with similar content.

Text Features Extraction based on TF-IDF Associating Semantic

An adaptive method for text domain similarity calculation

Research on Text Similarity Measurement Hybrid Algorithm with Term Semantic Information and TF-IDF Method

Improving Short Text Classification Through Better Feature Space Selection

TF-IDF Keyword Extraction Method Combining Context and Semantic Classification

Research of Chinese Text Classification Methods Based on Semantic Vector and Semantic Similarity

A Novel Text Clustering Algorithm Based on Inner Product Space Model of Semantic

Clustering Massive Text Data Streams by Semantic Smoothing Model

Research on Chinese Semantic Similarity Algorithm

The enhancement of TextRank algorithm by using word2vec and its application on topic extraction

A Study of Text Vectorization Method Combining Topic Model and Transfer Learning

Semantic similarity-aware feature selection and redundancy removal for text classification using joint mutual information

Clustering Text Data Streams

Semantic Similarity Computing Model Based on Multi Model Fine-Grained Nonlinear Fusion

Feature Dimension Reduction Short Text Clustering Combined with Semantic and Statistics

Garbage text classification filtering method Based on VSM

Text classification algorithm of tourist attractions subcategories with modified TF-IDF and Word2Vec

An Evaluation on Feature Selection for Text Clustering

A Study of the Application of Weight Distributing Method Combining Sentiment Dictionary and TF-IDF for Text Sentiment Analysis

Chinese Text Summarization Algorithm Based on Word2vec

Short Text Model Based on Strong Feature Thesaurus