Abstract:The explosive growth of digital video data renders a profound challenge to succinct, informative, and human-centric representations of video contents. This quickly-evolving research topic is typically called 'video abstraction'. We are motivated by the facts that the human brain is the end-evaluator of multimedia content and that the brain's responses can quantitatively reveal its attentional engagement in the comprehension of video. We propose a novel video abstraction paradigm which leverages functional magnetic resonance imaging (fMRI) to monitor and quantify the brain's responses to video stimuli. These responses are used to guide the extraction of visually informative segments from videos. Specifically, most relevant brain regions involved in video perception and cognition are identified to form brain networks. Then, the propensity for synchronization (PFS) derived from spectral graph theory is utilized over the brain networks to yield the benchmark attention curves based on the fMRI-measured brain responses to a number of training video streams. These benchmark attention curves are applied to guide and optimize the combinations of a variety of low-level visual features created by the Bayesian surprise model. In particular, in the training stage, the optimization objective is to ensure that the learned attentional model correlates well with the brain's responses and reflects the attention that viewers pay to video contents. In the application stage, the attention curves predicted by the learned and optimized attentional model serve as an effective benchmark to abstract testing videos. Evaluations on a set of video sequences from the TRECVID database demonstrate the effectiveness of the proposed framework.

Movie Fill in the Blank by Joint Learning from Video and Text with Adaptive Temporal Attention.

Movie Fill in the Blank with Adaptive Temporal Attention and Description Update.

Video Fill In the Blank using LR/RL LSTMs with Spatial-Temporal Attentions

ActionCLIP: Adapting Language-Image Pretrained Models for Video Action Recognition.

Adaptive Hierarchical Motion-Focused Model for Video Prediction.

MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Implicit Temporal Modeling with Learnable Alignment for Video Recognition

MovieLLM: Enhancing Long Video Understanding with AI-Generated Movies

Bidirectional Long-Short Term Memory for Video Description

Describing Video with Attention-Based Bidirectional LSTM

Blended Latent Diffusion under Attention Control for Real-World Video Editing

Alignment-guided Temporal Attention for Video Action Recognition

Video Abstraction Based on Fmri-Driven Visual Attention Model

Temporal Textual Localization in Video Via Adversarial Bi-Directional Interaction Networks

MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

FILS: Self-Supervised Video Feature Prediction In Semantic Language Space

FTAN: Exploring Frame-Text Attention for Lightweight Video Captioning.

Stand-Alone Inter-Frame Attention in Video Models

Enabling Language Models to Fill in the Blanks

FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

Frame Augmented Alternating Attention Network for Video Question Answering.