Label-Wise Document Pre-Training for Multi-Label Text Classification

Han Liu,Caixia Yuan,Xiaojie Wang
DOI: https://doi.org/10.1007/978-3-030-60450-9_51
2020-01-01
Abstract:A major challenge of multi-label text classification (MLTC) is to stimulatingly exploit possible label differences and label correlations. In this paper, we tackle this challenge by developing Label-Wise Pre-Training (LW-PT) method to get a document representation with label-aware information. The basic idea is that, a multi-label document can be represented as a combination of multiple label-wise representations, and that, correlated labels always cooccur in the same or similar documents. LW-PT implements this idea by constructing label-wise document classification tasks and trains label-wise document encoders. Finally, the pre-trained label-wise encoder is fine-tuned with the downstream MLTC task. Extensive experimental results validate that the proposed method has significant advantages over the previous state-of-the-art models and is able to discover reasonable label relationship. The code is released to facilitate other researchers.(https://github.com/laddie132/LW-PT).
What problem does this paper attempt to address?