Generative Paragraph Vector

Ruqing Zhang,Jiafeng Guo,Yanyan Lan,Jun Xu,Xueqi Cheng
DOI: https://doi.org/10.1007/978-3-030-01012-6_9
2018-01-01
Abstract:The recently introduced Paragraph Vector (PV) is an efficient method for learning high-quality distributed representations for texts. However, from the probabilistic view, PV is not a complete model since it only models the generation of words but not texts, leading to two major limitations. Firstly, without a text-level model, PV assumes the independence between texts and thus cannot leverage the corpus-wide information to help text representation learning. Secondly, without the generation model of texts, the inference of text representations outside of the training set becomes difficult. Although PV makes itself as an optimization problem so that one can obtain representations for new texts anyway, it loses the sound probabilistic interpretability in that way. To tackle these problems, we first introduce a Generative Paragraph Vector, an extension of the Distributed Bag of Words version of Paragraph Vector with a complete generative process. By defining the generation model over texts, we further incorporate text labels into the model and turn it into a supervised version, namely Supervised Generative Paragraph Vector. Experiments on five text classification benchmark collections show that both unsupervised and supervised model architectures can yield superior classification performance against the state-of-the-art counterparts.
What problem does this paper attempt to address?