Abstract:In this paper, we investigate the problem of abductive visual reasoning (AVR), which requires vision systems to infer the most plausible explanation for visual observations. Unlike previous work which performs visual reasoning on static images or synthesized scenes, we exploit long-term reasoning from instructional videos that contain a wealth of detailed information about the physical world. We conceptualize two tasks for this emerging and challenging topic. The primary task is AVR, which is based on the initial configuration and desired goal from an instructional video, and the model is expected to figure out what is the most plausible sequence of steps to achieve the goal. In order to avoid trivial solutions based on appearance information rather than reasoning, the second task called AVR++ is constructed, which requires the model to answer why the unselected options are less plausible. We introduce a new dataset called VideoABC, which consists of 46,354 unique steps derived from 11,827 instructional videos, formulated as 13,526 abductive reasoning questions with an average reasoning duration of 51 seconds. Through an adversarial hard hypothesis mining algorithm, non-trivial and high-quality problems are generated efficiently and effectively. To achieve human-level reasoning, we propose a Hierarchical Dual Reasoning Network (HDRNet) to capture the long-term dependencies among steps and observations. We establish a benchmark for abductive visual reasoning, and our method set state-of-the-arts on AVR ( similar to 74 %) and AVR++ ( similar to 45 %), and humans can easily achieve over 90% accuracy on these two tasks. The large performance gap reveals the limitation of current video understanding models on temporal reasoning and leaves substantial room for future research on this challenging problem. Our dataset and code are available at https://github.com/wl-zhao/VideoABC .

Explore Multi-Step Reasoning in Video Question Answering

Sequential Visual Reasoning for Visual Question Answering

MoReVQA: Exploring Modular Reasoning Models for Video Question Answering

Instance-sequence reasoning for video question answering

Video Question Answering with Semantic Disentanglement and Reasoning

Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering

Dynamic Spatio-Temporal Modular Network for Video Question Answering

Neural Reasoning, Fast and Slow, for Video Question Answering

From Representation to Reasoning: Towards both Evidence and Commonsense Reasoning for Video Question-Answering

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

DualVGR: A Dual-Visual Graph Reasoning Unit for Video Question Answering

Discovering Spatio-Temporal Rationales for Video Question Answering

Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering

Chain of Reasoning for Visual Question Answering.

VideoABC: A Real-World Video Dataset for Abductive Visual Reasoning

Event Graph Guided Compositional Spatial–Temporal Reasoning for Video Question Answering

TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering

Hierarchical Recurrent Contextual Attention Network for Video Question Answering

Video Question Answering for Surveillance

Joint Answering and Explanation for Visual Commonsense Reasoning

Video Question Answering: Datasets, Algorithms and Challenges