Abstract:Text classification is one of the fundamental tasks in natural language processing, which requires an agent to determine the most appropriate category for input sentences. Recently, deep neural networks have achieved impressive performance in this area, especially pretrained language models (PLMs). Usually, these methods concentrate on input sentences and corresponding semantic embedding generation. However, for another essential component: labels, most existing works either treat them as meaningless one-hot vectors or use vanilla embedding methods to learn label representations along with model training, underestimating the semantic information and guidance that these labels reveal. To alleviate this problem and better exploit label information, in this article, we employ self-supervised learning (SSL) in model learning process and design a novel self-supervised relation of relation (R <sup>2</sup> ) classification task for label utilization from a one-hot manner perspective. Then, we propose a novel () for text classification, in which text classification and R <sup>2</sup> classification are treated as optimization targets. Meanwhile, triplet loss is employed to enhance the analysis of differences and connections among labels. Moreover, considering that one-hot usage is still short of exploiting label information, we incorporate external knowledge from WordNet to obtain multiaspect descriptions for label semantic learning and extend to a novel () from a label embedding perspective. One step further, since these fine-grained descriptions may introduce unexpected noise, we develop a mutual interaction module to select appropriate parts from input sentences and labels simultaneously based on contrastive learning (CL) for noise mitigation. Extensive experiments on different text classification tasks reveal that can effectively improve the classification performance and can make better use of label information and further improve the performance. As a byproduct, we have released the codes to facilitate other research.

A Small-Sample Text Classification Model Based on Pseudo-Label Fusion Clustering Algorithm

Improving Short Text Classification Through Better Feature Space Selection

Multi-label Text Classification Model Based on Multi-level Constraint Augmentation and Label Association Attention

Description-Enhanced Label Embedding Contrastive Learning for Text Classification

Robust Representation Learning with Reliable Pseudo-labels Generation Via Self-Adaptive Optimal Transport for Short Text Clustering.

Short Text Classification of Chinese with Label Information Assisting

Category-wise Fine-Tuning: Resisting Incorrect Pseudo-Labels in Multi-Label Image Classification with Partial Labels

Image-text dual neural network with decision strategy for small-sample image classification

Integrated Image-Text Based on Semi-supervised Learning for Small Sample Instance Segmentation

Federated Learning for Short Text Clustering

Label-Specific Feature Augmentation for Long-Tailed Multi-Label Text Classification

CSAL: Self-adaptive Labeling based Clustering Integrating Supervised Learning on Unlabeled Data

Pseudo-Supervised Approach for Text Clustering Based on Consensus Analysis

A class-feature-centroid classifier for text categorization

Boosting Few-Shot Hyperspectral Image Classification Using Pseudo-Label Learning

Text Clustering as Classification with LLMs

Semi-Supervised Text Detection with Accurate Pseudo-Labels

Effective Multi-Label Active Learning for Text Classification

Less is More: Pseudo-Label Filtering for Continual Test-Time Adaptation

Text Grafting: Near-Distribution Weak Supervision for Minority Classes in Text Classification

Hierarchical Multi-label Text Classification: Self-adaption Semantic Awareness Network Integrating Text Topic and Label Level Information