Abstract:Video analysis is an important branch of computer vision due to its wide applications, ranging from video surveillance, video indexing, and retrieval to human computer interaction. All of the applications are based on a good video representation, which encodes video content into a feature vector with fixed length. Most existing methods treat video as a flat image sequence, but from our observations we argue that video is an information-intensive media with intrinsic hierarchical structure, which is largely ignored by previous approaches. Therefore, in this work, we represent the hierarchical structure of video with multiple granularities including, from short to long, single frame, consecutive frames (motion), short clip, and the entire video. Furthermore, we propose a novel deep learning framework to model each granularity individually. Specifically, we model the frame and motion granularities with 2D convolutional neural networks and model the clip and video granularities with 3D convolutional neural networks. Long Short-Term Memory networks are applied on the frame, motion, and clip to further exploit the long-term temporal clues. Consequently, the whole framework utilizes multi-stream CNNs to learn a hierarchical representation that captures spatial and temporal information of video. To validate its effectiveness in video analysis, we apply this video representation to action recognition task. We adopt a distribution-based fusion strategy to combine the decision scores from all the granularities, which are obtained by using a softmax layer on the top of each stream. We conduct extensive experiments on three action benchmarks (UCF101, HMDB51, and CCV) and achieve competitive performance against several state-of-the-art methods.

Multi-Level ResNets with Stacked SRUs for Action Recognition.

Multi-Level Recurrent Residual Networks for Action Recognition

Learning SpatioTemporal and Motion Features in a Unified 2D Network for Action Recognition

Revisiting the Spatial and Temporal Modeling for Few-shot Action Recognition

Action Recognition with Stacked Fisher Vectors.

Human Action Recognition Based on Three-Stream Network with Frame Sequence Features

Multi-scale residual network model combined with Global Average Pooling for action recognition

Action Recognition By Learning Deep Multi-Granular Spatio-Temporal Video Representation

Spatiotemporal Interaction Residual Networks with Pseudo3D for Video Action Recognition.

Action Recognition with Joint Attention on Multi-Level Deep Features

Mining Spatial and Spatio-Temporal ROIs for Action Recognition

Convolutional Neural Network-Based Video Super-Resolution for Action Recognition

A hybrid attention-guided ConvNeXt-GRU network for action recognition

Learning Hierarchical Video Representation for Action Recognition

Two-Stream Action Recognition-Oriented Video Super-Resolution.

StNet: Local and Global Spatial-Temporal Modeling for Action Recognition

Action-Stage Emphasized Spatiotemporal VLAD for Video Action Recognition

Sequential Segment Networks for Action Recognition

Unified Spatio-Temporal Attention Networks for Action Recognition in Videos.

Actor-Multi-Scale Context Bidirectional Higher Order Interactive Relation Network for Spatial-Temporal Action Localization

End-to-end Video-level Representation Learning for Action Recognition