Abstract:Video analysis is an important branch of computer vision due to its wide applications, ranging from video surveillance, video indexing, and retrieval to human computer interaction. All of the applications are based on a good video representation, which encodes video content into a feature vector with fixed length. Most existing methods treat video as a flat image sequence, but from our observations we argue that video is an information-intensive media with intrinsic hierarchical structure, which is largely ignored by previous approaches. Therefore, in this work, we represent the hierarchical structure of video with multiple granularities including, from short to long, single frame, consecutive frames (motion), short clip, and the entire video. Furthermore, we propose a novel deep learning framework to model each granularity individually. Specifically, we model the frame and motion granularities with 2D convolutional neural networks and model the clip and video granularities with 3D convolutional neural networks. Long Short-Term Memory networks are applied on the frame, motion, and clip to further exploit the long-term temporal clues. Consequently, the whole framework utilizes multi-stream CNNs to learn a hierarchical representation that captures spatial and temporal information of video. To validate its effectiveness in video analysis, we apply this video representation to action recognition task. We adopt a distribution-based fusion strategy to combine the decision scores from all the granularities, which are obtained by using a softmax layer on the top of each stream. We conduct extensive experiments on three action benchmarks (UCF101, HMDB51, and CCV) and achieve competitive performance against several state-of-the-art methods.

A Hierarchical Model of Shape and Appearance for Human Action Classification

A Hierarchical Pose-Based Approach to Complex Action Understanding Using Dictionaries of Actionlets and Motion Poselets

Online Robust Action Recognition Based on a Hierarchical Model

A Hierarchical Model for Human Action Recognition from Body-Parts

Reassessing Hierarchical Representation for Action Recognition in Still Images

A Hierarchical Model For Action Recognition Based On Body Parts

Human Action Recognition with Contextual Constraints Using a RGB-D Sensor

Action Recognition by Hierarchical Mid-level Action Elements

Extracting Hierarchical Spatial and Temporal Features for Human Action Recognition

Adaptive Hierarchical Motion-Focused Model for Video Prediction.

A Novel Hierarchical Framework for Human Action Recognition

Learning Hierarchical Video Representation for Action Recognition

Combining Sparse And Dense Descriptors With Temporal Semantic Structures For Robust Human Action Recognition

A Hierarchical Spatio-Temporal Model for Human Activity Recognition.

Shifting Perspective to See Difference: A Novel Multi-View Method for Skeleton Based Action Recognition

Human Action Recognition Based on Three-Stream Network with Frame Sequence Features

Embedding Motion and Structure Features for Action Recognition

Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions

A Matrix-Based Approach to Unsupervised Human Action Categorization

Video sketch: A middle-level representation for action recognition

Modeling Interactions Between Low-Level and High-Level Features for Human Action Recognition.