Answering Knowledge-Based Visual Questions Via the Exploration of Question Purpose

Lingyun Song,Jianao Li,Jun Liu,Yang,Xuequn Shang,Mingxuan Sun
DOI: https://doi.org/10.1016/j.patcog.2022.109015
IF: 8
2022-01-01
Pattern Recognition
Abstract:Visual question answering has been greatly advanced by deep learning technologies, but still remains an open problem subjected to two aspects of factors. First, previous works estimate the correctness of each candidate answer mainly by its semantic correlations with visual questions, overlooking the fact that some questions and their answers are semantically inconsistent. Second, previous works that require external knowledge mainly uses the knowledge facts retrieved by key words or visual objects. However, the retrieved knowledge facts may only be related to the semantics of the question, but are useless or even misleading for answer prediction. To address these issues, we investigate how to capture the pur-pose of visual questions and propose a Purpose Guided Visual Question Answering model, called PGVQA. It mainly has two appealing properties: (1) It can estimate the correctness of candidate answers based on the Question Purpose (QP) that reveals which aspects of the concept are examined by visual questions. This is helpful for avoiding the negative effect of the semantic inconsistency between answers and ques-tions. (2) It can incorporate the knowledge facts accordant with the QP into answer prediction, which helps to improve the probability of answering visual questions correctly. Empirical studies on benchmark datasets show that PGVQA achieves state-of-the-art performance.(c) 2022 Elsevier Ltd. All rights reserved.
What problem does this paper attempt to address?