Abstract:Text classification is one of the fundamental tasks in natural language processing, which requires an agent to determine the most appropriate category for input sentences. Recently, deep neural networks have achieved impressive performance in this area, especially pretrained language models (PLMs). Usually, these methods concentrate on input sentences and corresponding semantic embedding generation. However, for another essential component: labels, most existing works either treat them as meaningless one-hot vectors or use vanilla embedding methods to learn label representations along with model training, underestimating the semantic information and guidance that these labels reveal. To alleviate this problem and better exploit label information, in this article, we employ self-supervised learning (SSL) in model learning process and design a novel self-supervised relation of relation (R <sup>2</sup> ) classification task for label utilization from a one-hot manner perspective. Then, we propose a novel () for text classification, in which text classification and R <sup>2</sup> classification are treated as optimization targets. Meanwhile, triplet loss is employed to enhance the analysis of differences and connections among labels. Moreover, considering that one-hot usage is still short of exploiting label information, we incorporate external knowledge from WordNet to obtain multiaspect descriptions for label semantic learning and extend to a novel () from a label embedding perspective. One step further, since these fine-grained descriptions may introduce unexpected noise, we develop a mutual interaction module to select appropriate parts from input sentences and labels simultaneously based on contrastive learning (CL) for noise mitigation. Extensive experiments on different text classification tasks reveal that can effectively improve the classification performance and can make better use of label information and further improve the performance. As a byproduct, we have released the codes to facilitate other research.

Unsupervised multimodal learning for image-text relation classification in tweets

Modality-invariant Temporal Representation Learning for Multimodal Sentiment Classification

Semi-Supervised Dual Relation Learning for Multi-Label Classification

Detection of Illicit Drug Trafficking Events on Instagram: A Deep Multimodal Multilabel Learning Approach

Improving Multimodal Classification of Social Media Posts by Leveraging Image-Text Auxiliary Tasks

Social Image-text Sentiment Classification With Cross-Modal Consistency and Knowledge Distillation

RIVA: A Pre-trained Tweet Multimodal Model Based on Text-image Relation for Multimodal NER.

A multi‐label social short text classification method based on contrastive learning and improved ml‐KNN

I2SRM: Intra- and Inter-Sample Relationship Modeling for Multimodal Information Extraction

Graph-based Multimodal Semi-Supervised Image Classification

Joint Intermodal and Intramodal Label Transfers for Extremely Rare or Unseen Classes

Contrastive Graph Multimodal Model for Text Classification in Videos

Social Image Sentiment Analysis by Exploiting Multimodal Content and Heterogeneous Relations

Multimodal Sentiment Analysis With Image-Text Interaction Network

Image-text Retrieval: A Survey on Recent Research and Development

TCGM: an Information-Theoretic Framework for Semi-Supervised Multi-Modality Learning

A multimodal sentiment recognition method based on attention mechanism

Description-Enhanced Label Embedding Contrastive Learning for Text Classification

Borrowing Human Senses: Comment-Aware Self-Training for Social Media Multimodal Classification

Understanding, Categorizing and Predicting Semantic Image-Text Relations

RoMo: Robust Unsupervised Multimodal Learning with Noisy Pseudo Labels