Abstract:AbstractVisual Question Answering (VQA) is a challenging task that has gained increasing attention from both the computer vision and the natural language processing communities in recent years. Given a question in natural language, a VQA system is designed to automatically generate the answer according to the referenced visual content. Though there recently has been much intereset in this topic, the existing work of visual question answering mainly focuses on a single static image, which is only a small part of the dynamic and sequential visual data in the real world. As a natural extension, video question answering (VideoQA) is less explored. Because of the inherent temporal structure in the video, the approaches of ImageQA may be ineffectively applied to video question answering. In this article, we not only take the spatial and temporal dimension of video content into account but also employ an external knowledge base to improve the answering ability of the network. More specifically, we propose a knowledge-based progressive spatial-temporal attention network to tackle this problem. We obtain both objects and region features of the video frames from a region proposal network. The knowledge representation is generated by a word-level attention mechanism using the comment information of each object that is extracted from DBpedia. Then, we develop a question-knowledge-guided progressive spatial-temporal attention network to learn the joint video representation for video question answering task. We construct a large-scale video question answering dataset. The extensive experiments based on two different datasets validate the effectiveness of our method.

Long-Term Video Question Answering Via Multimodal Hierarchical Memory Attentive Networks

Video Question Answering Via Hierarchical Dual-Level Attention Network Learning.

Memory Augmented Deep Recurrent Neural Network for Video Question Answering

Heterogeneous Memory Enhanced Multimodal Attention Model for Video Question Answering

Video Question Answering Via Multi-Granularity Temporal Attention Network Learning

Multi-Turn Video Question Answering Via Multi-Stream Hierarchical Attention Context Network

Video Question Answering Via Grounded Cross-Attention Network Learning.

Multi-Turn Video Question Answering via Hierarchical Attention Context Reinforced Networks

Hierarchical Temporal Fusion of Multi-grained Attention Features for Video Question Answering

Multilevel Hierarchical Network with Multiscale Sampling for Video Question Answering

Video Question Answering via Attribute-Augmented Attention Network Learning

Video Question Answering Via Hierarchical Spatio-Temporal Attention Networks

Hierarchical Recurrent Contextual Attention Network for Video Question Answering

Advancing Video Question Answering with a Multi-modal and Multi-layer Question Enhancement Network

Modular Blended Attention Network for Video Question Answering

Frame Augmented Alternating Attention Network for Video Question Answering.

Long-Form Video Question Answering via Dynamic Hierarchical Reinforced Networks

Video Question Answering via Knowledge-based Progressive Spatial-Temporal Attention Network

MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Video Question Answering Using a Forget Memory Network

A Better Way to Attend: Attention with Trees for Video Question Answering