Abstract:We first propose a new low-level visual feature, called spatio-temporal context distribution feature of interest points, to describe human actions. Each action video is expressed as a set of relative XYT coordinates between pairwise interest points in a local region. We learn a global Gaussian mixture model (GMM) (referred to as a universal background model) using the relative coordinate features from all the training videos, and then we represent each video as the normalized parameters of a video-specific GMM adapted from the global GMM. In order to capture the spatio-temporal relationships at different levels, multiple GMMs are utilized to describe the context distributions of interest points over multiscale local regions. Motivated by the observation that some actions share similar motion patterns, we additionally propose a novel mid-level class correlation feature to capture the semantic correlations between different action classes. Each input action video is represented by a set of decision values obtained from the pre-learned classifiers of all the action classes, with each decision value measuring the likelihood that the input video belongs to the corresponding action class. Moreover, human actions are often associated with some specific natural environments and also exhibit high correlation with particular scene classes. It is therefore beneficial to utilize the contextual scene information for action recognition. In this paper, we build the high-level co-occurrence relationship between action classes and scene classes to discover the mutual contextual constraints between action and scene. By treating the scene class label as a latent variable, we propose to use the latent structural SVM (LSSVM) model to jointly capture the compatibility between multilevel action features (e.g., low-level visual context distribution feature and the corresponding mid-level class correlation feature) and action classes, the compatibility between multilevel scene features (i.e., SIFT feature and the corresponding class correlation feature) and scene classes, and the contextual relationship between action classes and scene classes. Extensive experiments on UCF Sports, YouTube and UCF50 datasets demonstrate the effectiveness of the proposed multilevel features and action-scene interaction based LSSVM model for human action recognition. Moreover, our method generally achieves higher recognition accuracy than other state-of-the-art methods on these datasets.

Weakly Supervised Action Recognition And Localization Using Web Images

Transfer Latent SVM for Joint Recognition and Localization of Actions in Videos.

Weakly-Supervised Action Localization by Hierarchically-structured Latent Attention Modeling

Weakly-Supervised Action Recognition and Localization via Knowledge Transfer.

Weakly-Supervised Action Localization by Hierarchical Attention Mechanism with Multi-Scale Fusion Strategies

Modeling Sub-Actions for Weakly Supervised Temporal Action Localization

Learning Transferable Self-attentive Representations for Action Recognition in Untrimmed Videos with Weak Supervision

TwinNet: Twin Structured Knowledge Transfer Network for Weakly Supervised Action Localization

Weakly-supervised action localization via embedding-modeling iterative optimization

Weakly-Supervised Action Localization by Generative Attention Modeling

Weakly-Supervised Temporal Action Localization Based on Attention Regularization

Weakly-supervised Action Localization Via Hierarchical Mining.

Weakly-Supervised Temporal Action Localization with Regional Similarity Consistency

Perceiving Local Relative Motion and Global Correlations for Weakly Supervised Group Activity Recognition.

Action Recognition Using Multilevel Features and Latent Structural SVM

Weakly-Supervised Temporal Action Localization Via Cross-Stream Collaborative Learning.

Spatio-Temporal Action Localization in a Weakly Supervised Setting

Action-Semantic Consistent Knowledge for Weakly-Supervised Action Localization

Adaptive Mutual Supervision for Weakly-Supervised Temporal Action Localization

Cross-Video Contextual Knowledge Exploration and Exploitation for Ambiguity Reduction in Weakly Supervised Temporal Action Localization

Exploring Sub-Action Granularity for Weakly Supervised Temporal Action Localization