Abstract:Text classification is one of the fundamental tasks in natural language processing, which requires an agent to determine the most appropriate category for input sentences. Recently, deep neural networks have achieved impressive performance in this area, especially pretrained language models (PLMs). Usually, these methods concentrate on input sentences and corresponding semantic embedding generation. However, for another essential component: labels, most existing works either treat them as meaningless one-hot vectors or use vanilla embedding methods to learn label representations along with model training, underestimating the semantic information and guidance that these labels reveal. To alleviate this problem and better exploit label information, in this article, we employ self-supervised learning (SSL) in model learning process and design a novel self-supervised relation of relation (R <sup>2</sup> ) classification task for label utilization from a one-hot manner perspective. Then, we propose a novel () for text classification, in which text classification and R <sup>2</sup> classification are treated as optimization targets. Meanwhile, triplet loss is employed to enhance the analysis of differences and connections among labels. Moreover, considering that one-hot usage is still short of exploiting label information, we incorporate external knowledge from WordNet to obtain multiaspect descriptions for label semantic learning and extend to a novel () from a label embedding perspective. One step further, since these fine-grained descriptions may introduce unexpected noise, we develop a mutual interaction module to select appropriate parts from input sentences and labels simultaneously based on contrastive learning (CL) for noise mitigation. Extensive experiments on different text classification tasks reveal that can effectively improve the classification performance and can make better use of label information and further improve the performance. As a byproduct, we have released the codes to facilitate other research.

Label Distribution Learning-Enhanced Dual-KNN for Text Classification

RW.KNN: a proposed random walk KNN algorithm for multi-label classification.

A Bayesian Network nearest k-labels method for Multi-label classification

Large Margin Weighted K-Nearest Neighbors Label Distribution Learning for Classification.

A Debiased Nearest Neighbors Framework for Multi-Label Text Classification

Label Distribution Learning by Exploiting Feature-Label Correlations Locally

Contrastive Learning-Enhanced Nearest Neighbor Mechanism for Multi-Label Text Classification.

$k$-Nearest Neighbor Augmented Neural Networks for Text Classification

An improved ML-kNN approach for multi-label text categorization

Residual K-Nearest Neighbors Label Distribution Learning

Gate-Attention and Dual-End Enhancement Mechanism for Multi-Label Text Classification

Learning Neural Networks for Text Classification by Exploiting Label Relations

Neighbor selection for multilabel classification

Probability Knowledge Acquisition from Unlabeled Instance Based on Dual Learning

A Dual-Branch Learning Model with Gradient-Balanced Loss for Long-Tailed Multi-Label Text Classification

Neural Text Classification by Jointly Learning to Cluster and Align

TK-KNN: A Balanced Distance-Based Pseudo Labeling Approach for Semi-Supervised Intent Classification

An Improved Knn Text Classification Algorithm Based On Density

Nearest Neighbor Method Based on Local Distribution for Classification.

Description-Enhanced Label Embedding Contrastive Learning for Text Classification

A Dual-CNN Model for Multi-label Classification by Leveraging Co-occurrence Dependencies Between Labels.