Identifying new categories in community question answering archives: a topic modeling approach.

Yajie Miao,Chunping Li,Jie Tang,Lili Zhao
DOI: https://doi.org/10.1145/1871437.1871701
2010-01-01
Abstract:Community Question Answering (CQA) services have evolved into a popular way of information seeking and providing. User-posted questions in CQA are generally organized into hierarchical categories. In this paper, we define and study a novel problem which is referred to as New Category Identification (NCI) in CQA question archives. New Category Identification is primarily concerned with detecting and characterizing new or emerging categories which are not included in the existing category hierarchy. We define this problem formally, and propose both unsupervised and semi-supervised topic modeling methods to solve it. Experiments with a ground-truth set built from Yahoo! Answers show that our methods identify and interpret new categories effectively.
What problem does this paper attempt to address?