Abstract:Visual question answering (VQA) is challenging, because it requires a simultaneous understanding of both visual content of images and textual content of questions. To support the VQA task, we need to find good solutions for the following three issues: 1) fine-grained feature representations for both the image and the question; 2) multimodal feature fusion that is able to capture the complex interactions between multimodal features; and 3) automatic answer prediction that is able to consider the complex correlations between multiple diverse answers for the same question. For fine-grained image and question representations, a "coattention" mechanism is developed using a deep neural network (DNN) architecture to jointly learn the attentions for both the image and the question, which can allow us to reduce the irrelevant features effectively and obtain more discriminative features for image and question representations. For multimodal feature fusion, a generalized multimodal factorized high-order pooling approach (MFH) is developed to achieve more effective fusion of multimodal features by exploiting their correlations sufficiently, which can further result in superior VQA performance as compared with the state-of-the-art approaches. For answer prediction, the Kullback-Leibler divergence is used as the loss function to achieve precise characterization of the complex correlations between multiple diverse answers with the same or similar meaning, which can allow us to achieve faster convergence rate and obtain slightly better accuracy on answer prediction. A DNN architecture is designed to integrate all these aforementioned modules into a unified model for achieving superior VQA performance. With an ensemble of our MFH models, we achieve the state-of-the-art performance on the large-scale VQA data sets and win the runner-up in VQA Challenge 2017.

MUTAN: Multimodal Tucker Fusion for Visual Question Answering

Dual Path Multi-Modal High-Order Features for Textual Content Based Visual Question Answering

Simple and Effective Visual Question Answering in a Single Modality

MIRTT: Learning Multimodal Interaction Representations from Trilinear Transformers for Visual Question Answering

Compact Tensor Pooling for Visual Question Answering

Multimodal Integration of Human-Like Attention in Visual Question Answering

Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering

Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Temporal Pyramid Transformer with Multimodal Interaction for Video Question Answering

Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion

Generalized Hadamard-Product Fusion Operators for Visual Question Answering

Video Pivoting Unsupervised Multi-Modal Machine Translation

DecomVQANet: Decomposing visual question answering deep network via tensor decomposition and regression

Improved Fusion of Visual and Language Representations by Dense Symmetric Co-Attention for Visual Question Answering

Collaborative Modality Fusion for Mitigating Language Bias in Visual Question Answering

Multi-Modal Fusion Transformer for Visual Question Answering in Remote Sensing

Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA

Multi-modal Factorized Bilinear Pooling with Co-attention Learning for Visual Question Answering.

Multi-stage hybrid embedding fusion network for visual question answering

Multimodal Bi-direction Guided Attention Networks for Visual Question Answering

Bridging the Gap Between 2D and 3D Visual Question Answering: A Fusion Approach for 3D VQA