Incorporating Social Context And Domain Knowledge For Entity Recognition

Jie Tang,Zhanpeng Fang,Jimeng Sun
DOI: https://doi.org/10.1145/2736277.2741135
2015-01-01
Abstract:Recognizing entity instances in documents according to a knowledge base is a fundamental problem in many data mining applications. The problem is extremely challenging for short documents in complex domains such as social media and biomedical domains. Large concept spaces and instance ambiguity are key issues that need to be addressed.Most of the documents are created in a social context by common authors via social interactions, such as reply and citations. Such social contexts are largely ignored in the instance-recognition literature. How can users' interactions help entity instance recognition? How can the social context be modeled so as to resolve the ambiguity of different instances?In this paper, we propose the SOCINST model to formalize the problem into a probabilistic model. Given a set of short documents (e.g., tweets or paper abstracts) posted by users who may connect with each other, SOCINST can automatically construct a context of subtopics for each instance, with each subtopic representing one possible meaning of the instance. The model is also able to incorporate social relationships between users to help build social context. We further incorporate domain knowledge into the model using a Dirichlet tree distribution.We evaluate the proposed model on three different genres of datasets: ICDM' 12 Contest, Weibo, and I2B2. In ICDM' 12 Contest, the proposed model clearly outperforms (+21.4%; p << 1e - 5 with t-test) all the top contestants. In Weibo and I2B2, our results also show that the recognition accuracy of SOCINST is up to 5.3-26.6% better than those of several alternative methods.
What problem does this paper attempt to address?