Abstract:Text based Visual Question Answering (TextVQA) is a recently raised challenge requiring models to read text in images and answer natural language questions by jointly reasoning over the question, textual information and visual content. Introduction of this new modality - Optical Character Recognition (OCR) tokens ushers in demanding reasoning requirements. Most of the state-of-the-art (SoTA) VQA methods fail when answer these questions because of three reasons: (1) poor text reading ability; (2) lack of textual-visual reasoning capacity; and (3) choosing discriminative answering mechanism over generative couterpart (although this has been further addressed by M4C). In this paper, we propose an end-to-end structured multimodal attention (SMA) neural network to mainly solve the first two issues above. SMA first uses a structural graph representation to encode the object-object, object-text and text-text relationships appearing in the image, and then designs a multimodal graph attention network to reason over it. Finally, the outputs from the above modules are processed by a global-local attentional answering module to produce an answer splicing together tokens from both OCR and general vocabulary iteratively by following M4C. Our proposed model outperforms the SoTA models on TextVQA dataset and two tasks of ST-VQA dataset among all models except pre-training based TAP. Demonstrating strong reasoning ability, it also won first place in TextVQA Challenge 2020. We extensively test different OCR methods on several reasoning models and investigate the impact of gradually increased OCR performance on TextVQA benchmark. With better OCR results, different models share dramatic improvement over the VQA accuracy, but our model benefits most blessed by strong textual-visual reasoning ability. To grant our method an upper bound and make a fair testing base available for further works, we also provide human-annotated ground-truth OCR annotations for the TextVQA dataset, which were not given in the original release. The code and ground-truth OCR annotations for the TextVQA dataset are available at https://github.com/ChenyuGAO-CS/SMA.

MAPM: multiscale attention pre-training model for TextVQA

Dual Path Multi-Modal High-Order Features for Textual Content Based Visual Question Answering

Two-Stage Multimodality Fusion for High-Performance Text-Based Visual Question Answering.

Multi-Modal Learning with Text Merging for TEXTVQA.

Cross-Modal Self-Attention with Multi-Task Pre-Training for Medical Visual Question Answering

Masked Vision and Language Pre-training with Unimodal and Multimodal Contrastive Losses for Medical Visual Question Answering

Multiscale Feature Extraction and Fusion of Image and Text in VQA

MoCA: Incorporating Multi-stage Domain Pretraining and Cross-guided Multimodal Attention for Textbook Question Answering

A multimodal attention fusion network with a dynamic vocabulary for TextVQA

Structured Multimodal Attentions for TextVQA

MoCA: Incorporating domain pretraining and cross attention for textbook question answering

MGA-VQA: Multi-Granularity Alignment for Visual Question Answering

CrossMAE: Cross Modality Masked Autoencoders for Region-Aware Audio-Visual Pretraining

RAMM: Retrieval-augmented Biomedical Visual Question Answering with Multi-modal Pre-training

MISS: A Generative Pretraining and Finetuning Approach for Med-VQA

Pre-training Model Based on Parallel Cross-Modality Fusion Layer

Multi-task Paired Masking with Alignment Modeling for Medical Vision-Language Pre-training

Multitask Learning for Visual Question Answering

AMAM: An Attention-based Multimodal Alignment Model for Medical Visual Question Answering

Enhancing visual question answering with a two‐way co‐attention mechanism and integrated multimodal features