Finding a good query-related topic for boosting pseudo-relevance feedback

Zheng Ye,Jimmy Xiangji Huang,Hongfei Lin
DOI: https://doi.org/10.1002/asi.21501
2011-01-01
Journal of the American Society for Information Science and Technology
Abstract:Pseudo-relevance feedback (PRF) via query expansion (QE) assumes that the top-ranked documents from the first-pass retrieval are relevant. The most informative terms in the pseudo-relevant feedback documents are then used to update the original query representation in order to boost the retrieval performance. Most current PRF approaches estimate the importance of the candidate expansion terms based on their statistics on document level. However, a document for PRF may consist of different topics, which may not be all related to the query even if the document is judged relevant. The main argument of this article is the proposal to conduct PRF on a granularity smaller than on the document level. In this article, we propose a topic-based feedback model with three different strategies for finding a good query-related topic based on the Latent Dirichlet Allocation model. The experimental results on four representative TREC collections show that QE based on the derived topic achieves statistically significant improvements over a strong feedback model in the language modeling framework, which updates the query representation based on the top-ranked documents. © 2011 Wiley Periodicals, Inc.
What problem does this paper attempt to address?