Abstract:Multi-label text classification tasks face challenges such as sample diversity, complexity, and the need for effective utilization of label correlations. In this paper, we propose a model that integrates multi-granularity fusion of text sequence features and label semantic correlation information. Our model leverages graph convolutional networks to extract label semantic correlation, which enhances classification performance for samples with similar labels and addresses label omission issues. Additionally, text convolutional neural networks are employed to extract multi-granularity sense group features from text sequences, calculate their similarity with semantic correlation label distributions, and dynamically adjust the similarity between text context and label information. This approach tackles the limitations of feature extraction in short texts and label confusion. We replace the original multi-hot label encoding in model training with a label distribution that fuses text multi-granularity sense group features and label correlation information, using a more precise encoding method for soft alignment based on label probability distributions. This enhances the model's resilience to noisy data, avoiding the issue of assigning high-confidence probabilities to incorrect categories due to hard-coded supervision. Our model's performance improvement on noisy datasets significantly surpasses that achieved by label smoothing. Extensive experiments on three legal text datasets and two generalized multi-label datasets demonstrate the model's excellent performance. Our approach is applicable in various real-world scenarios, such as legal judgment prediction, news categorization, and recommendation systems, where accurate multi-label classification is crucial. Ablation and experiments on noisy datasets validate the model's effectiveness and robustness.

Label correlation mixture model for multi-label text categorization

Label Correlation Mixture Model: A Supervised Generative Approach to Multilabel Spoken Document Categorization

Multi-label Text Classification Based on the Label Correlation Mixture Model.

A Generative Probabilistic Model for Multi-label Classification

Muli-label Text Categorization with Hidden Components.

Minimum classification error rate training of supervised topic mixture model for multi-label text categorization

LCBM: A Multi-View Probabilistic Model for Multi-Label Classification

Capturing Correlations of Multiple Labels: A Generative Probabilistic Model for Multi-Label Learning

Research of multi-label text classification based on label attention and correlation networks

Multi-label classification by exploiting label correlations

MFLSCI: Multi-granularity fusion and label semantic correlation information for multi-label legal text classification

Partial Multi-label Learning with Label and Feature Collaboration

Generative Multi-Label Correlation Learning

Correlation concept-cognitive learning model for multi-label classification

Enhancing Label Correlation Feedback in Multi-Label Text Classification via Multi-Task Learning

Multi-Label Learning by Exploiting Label Correlations with LDA

A Label Information Aware Model for Multi-label Text Classification

Multi-label Text Categorization with Joint Learning Predictions-as-Features Method

CorrLog: Correlated Logistic Models for Joint Prediction of Multiple Labels.

Multi-label Text Classification Model Based on Multi-level Constraint Augmentation and Label Association Attention

Rethinking Modal-oriented Label Correlations for Multi-modal Multi-label Learning