Abstract:Recent studies have shown promising results on utilizing large pre-trained image-language models for video question answering. While these image-language models can efficiently bootstrap the representation learning of video-language models, they typically concatenate uniformly sampled video frames as visual inputs without explicit language-aware, temporal modeling. When only a portion of a video input is relevant to the language query, such uniform frame sampling can often lead to missing important visual cues. Although humans often find a video moment to focus on and rewind the moment to answer questions, training a query-aware video moment localizer often requires expensive annotations and high computational costs. To address this issue, we propose Self-Chained Video Localization-Answering (SeViLA), a novel framework that leverages a single image-language model (BLIP-2) to tackle both temporal keyframe localization and QA on videos. SeViLA framework consists of two modules: Localizer and Answerer, where both are parameter-efficiently fine-tuned from BLIP-2. We propose two ways of chaining these modules for cascaded inference and self-refinement. First, in the forward chain, the Localizer finds multiple language-aware keyframes in a video, which the Answerer uses to predict the answer. Second, in the reverse chain, the Answerer generates keyframe pseudo-labels to refine the Localizer, alleviating the need for expensive video moment localization annotations. Our SeViLA framework outperforms several strong baselines on 5 challenging video QA and event prediction benchmarks, and achieves the state-of-the-art in both fine-tuning (NExT-QA, STAR) and zero-shot (NExT-QA, STAR, How2QA, VLEP) settings. We also analyze the impact of Localizer, comparisons of Localizer with other temporal localization models, pre-training/self-refinement of Localizer, and varying the number of keyframes.

Language-aware Visual Semantic Distillation for Video Question Answering

VideoDistill: Language-aware Vision Distillation for Video Question Answering

Simple and Effective Visual Question Answering in a Single Modality

Overcoming Language Priors In Vqa Via Decomposed Linguistic Representations

Towards Developing a Multilingual and Code-Mixed Visual Question Answering System by Knowledge Distillation

Distilled Dual-Encoder Model for Vision-Language Understanding

Adaptive Spatio-Temporal Graph Enhanced Vision-Language Representation for Video QA

Equivariant and Invariant Grounding for Video Question Answering

LiVLR: A Lightweight Visual-Linguistic Reasoning Framework for Video Question Answering

Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering

ViLA: Efficient Video-Language Alignment for Video Question Answering

Unlocking Video-LLM via Agent-of-Thoughts Distillation

Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models

Answering Diverse Questions via Text Attached with Key Audio-Visual Clues

Enhancing Continual Learning in Visual Question Answering with Modality-Aware Feature Distillation

Video Question Answering Via Gradually Refined Attention over Appearance and Motion

Distilling Vision-Language Models on Millions of Videos

Self-Chained Image-Language Model for Video Localization and Question Answering

Boosting Audio Visual Question Answering via Key Semantic-Aware Cues

Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering

Video Question Answering Using CLIP-Guided Visual-Text Attention