Aligned Dual Channel Graph Convolutional Network for Visual Question Answering.

Qingbao Huang,Jielong Wei,Yi Cai,Changmeng Zheng,Junying Chen,Ho-fung Leung,Qing Li
DOI: https://doi.org/10.18653/v1/2020.acl-main.642
2020-01-01
Abstract:Visual question answering aims to answer the natural language question about a given image. Existing graph-based methods only focus on the relations between objects in an image and neglect the importance of the syntactic dependency relations between words in a question. To simultaneously capture the relations between objects in an image and the syntactic dependency relations between words in a question, we propose a novel dual channel graph convolutional network (DC-GCN) for better combining visual and textual advantages. The DC-GCN model consists of three parts: an I-GCN module to capture the relations between objects in an image, a Q-GCN module to capture the syntactic dependency relations between words in a question, and an attention alignment module to align image representations and question representations. Experimental results show that our model achieves comparable performance with the state-of-the-art approaches.
What problem does this paper attempt to address?