Abstract:Video analysis is an important branch of computer vision due to its wide applications, ranging from video surveillance, video indexing, and retrieval to human computer interaction. All of the applications are based on a good video representation, which encodes video content into a feature vector with fixed length. Most existing methods treat video as a flat image sequence, but from our observations we argue that video is an information-intensive media with intrinsic hierarchical structure, which is largely ignored by previous approaches. Therefore, in this work, we represent the hierarchical structure of video with multiple granularities including, from short to long, single frame, consecutive frames (motion), short clip, and the entire video. Furthermore, we propose a novel deep learning framework to model each granularity individually. Specifically, we model the frame and motion granularities with 2D convolutional neural networks and model the clip and video granularities with 3D convolutional neural networks. Long Short-Term Memory networks are applied on the frame, motion, and clip to further exploit the long-term temporal clues. Consequently, the whole framework utilizes multi-stream CNNs to learn a hierarchical representation that captures spatial and temporal information of video. To validate its effectiveness in video analysis, we apply this video representation to action recognition task. We adopt a distribution-based fusion strategy to combine the decision scores from all the granularities, which are obtained by using a softmax layer on the top of each stream. We conduct extensive experiments on three action benchmarks (UCF101, HMDB51, and CCV) and achieve competitive performance against several state-of-the-art methods.

Action Recognition with Stacked Fisher Vectors.

Hyper-Fisher Vectors for Action Recognition

Good Practices for Learning to Recognize Actions Using FV and VLAD

Action-Stage Emphasized Spatiotemporal VLAD for Video Action Recognition

Action Recognition By Learning Deep Multi-Granular Spatio-Temporal Video Representation

Human Action Recognition Based on Three-Stream Network with Frame Sequence Features

Reassessing Hierarchical Representation for Action Recognition in Still Images

Learning Hierarchical Video Representation for Action Recognition

Multi-view Super Vector for Action Recognition

3D Action Recognition Using Multi-Temporal Depth Motion Maps and Fisher Vector

MoFAP: A Multi-level Representation for Action Recognition

DA-VLAD: Discriminative Action Vector of Locally Aggregated Descriptors for Action Recognition

Discriminative Multi-View Subspace Feature Learning for Action Recognition

Hybrid super vector with improved dense trajectories for action recognition

F2D-SIFPNet: a Frequency 2D Slow-I-Fast-P Network for Faster Compressed Video Action Recognition

A Joint Evaluation of Dictionary Learning and Feature Encoding for Action Recognition.

3-Stream Convolutional Networks for Video Action Recognition with Hybrid Motion Field

Spatio-temporal stacking model for skeleton-based action recognition

ActionVLAD: Learning Spatio-Temporal Aggregation for Action Classification

Action recognition and detection by combining motion and appearance features

End-to-end Video-level Representation Learning for Action Recognition