Abstract:Text classification is one of the fundamental tasks in natural language processing, which requires an agent to determine the most appropriate category for input sentences. Recently, deep neural networks have achieved impressive performance in this area, especially pretrained language models (PLMs). Usually, these methods concentrate on input sentences and corresponding semantic embedding generation. However, for another essential component: labels, most existing works either treat them as meaningless one-hot vectors or use vanilla embedding methods to learn label representations along with model training, underestimating the semantic information and guidance that these labels reveal. To alleviate this problem and better exploit label information, in this article, we employ self-supervised learning (SSL) in model learning process and design a novel self-supervised relation of relation (R <sup>2</sup> ) classification task for label utilization from a one-hot manner perspective. Then, we propose a novel () for text classification, in which text classification and R <sup>2</sup> classification are treated as optimization targets. Meanwhile, triplet loss is employed to enhance the analysis of differences and connections among labels. Moreover, considering that one-hot usage is still short of exploiting label information, we incorporate external knowledge from WordNet to obtain multiaspect descriptions for label semantic learning and extend to a novel () from a label embedding perspective. One step further, since these fine-grained descriptions may introduce unexpected noise, we develop a mutual interaction module to select appropriate parts from input sentences and labels simultaneously based on contrastive learning (CL) for noise mitigation. Extensive experiments on different text classification tasks reveal that can effectively improve the classification performance and can make better use of label information and further improve the performance. As a byproduct, we have released the codes to facilitate other research.

Word-class embeddings for multiclass text classification

Enhanced Double-Carrier Word Embedding Via Phonetics and Writing

Improve Word Embedding Using Both Writing and Pronunciation.

Word Embeddings and Their Use In Sentence Classification Tasks

Utility of General and Specific Word Embeddings for Classifying Translational Stages of Research

Bag-of-Embeddings for Text Classification.

Category Enhanced Word Embedding.

Expanding the Text Classification Toolbox with Cross-Lingual Embeddings

Pre-Trained Multi-View Word Embedding Using Two-Side Neural Network

An Exploration Of Semantic Relations In Neural Word Embeddings Using Extrinsic Knowledge

Boosting Semantic Segmentation from the Perspective of Explicit Class Embeddings

A Privacy-Preserving Word Embedding Text Classification Model Based on Privacy Boundary Constructed by Deep Belief Network

Word Embedding for Text Classification: Efficient CNN and Bi-GRU Fusion Multi Attention Mechanism

Unified Embedding: Battle-Tested Feature Representations for Web-Scale ML Systems

Description-Enhanced Label Embedding Contrastive Learning for Text Classification

Learning Effective Word Embedding Using Morphological Word Similarity

Leveraging Semantic Segmentation Masks with Embeddings for Fine-Grained Form Classification

End-to-End Text Classification via Image-based Embedding using Character-level Networks

On Embedding Implementations in Text Ranking and Classification Employing Graphs

Estimating Text Similarity based on Semantic Concept Embeddings

Deep learning with word embeddings improves biomedical named entity recognition