Abstract:One significant factor we expect the video representation learning to capture, especially in contrast with the image representation learning, is the object motion. However, we found that in the current mainstream video datasets, some action categories are highly related with the scene where the action happens, making the model tend to degrade to a solution where only the scene information is encoded. For example, a trained model may predict a video as playing football simply because it sees the field, neglecting that the subject is dancing as a cheerleader on the field. This is against our original intention towards the video representation learning and may bring scene bias on a different dataset that can not be ignored. In order to tackle this problem, we propose to decouple the scene and the motion (DSM) with two simple operations, so that the model attention towards the motion information is better paid. Specifically, we construct a positive clip and a negative clip for each video. Compared to the original video, the positive/negative is motion-untouched/broken but scene-broken/untouched by Spatial Local Disturbance and Temporal Local Disturbance. Our objective is to pull the positive closer while pushing the negative farther to the original clip in the latent space. In this way, the impact of the scene is weakened while the temporal sensitivity of the network is further enhanced. We conduct experiments on two tasks with various backbones and different pre-training datasets, and find that our method surpass the SOTA methods with a remarkable 8.1% and 8.8% improvement towards action recognition task on the UCF101 and HMDB51 datasets respectively using the same backbone.

Enhancing Motion Visual Cues for Self-Supervised Video Representation Learning

Point Contrastive Prediction with Semantic Clustering for Self-Supervised Learning on Point Cloud Videos

Self-Supervised Video Representation Learning with Motion-Contrastive Perception

Masked Motion Encoding for Self-Supervised Video Representation Learning

Cross-view motion consistent self-supervised video inter-intra contrastive for action representation understanding

Mitigating background bias in self-supervised video representation learning

Motion Sensitive Contrastive Learning for Self-supervised Video Representation

Enhancing Unsupervised Video Representation Learning by Decoupling the Scene and the Motion

Self-supervised Spatio-temporal Representation Learning for Videos by Predicting Motion and Appearance Statistics

Motion-Focused Contrastive Learning of Video Representations*

Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization

Self-supervised pretext task collaborative multi-view contrastive learning for video action recognition

Temporally-Embedded Self-Supervised Video Representation Learning

Self-Supervised Video Representation Learning in a Heuristic Decoupled Perspective

Self-supervised Video Representation Learning Using Inter-intra Contrastive Framework

Self-supervised Video Representation Learning via Capturing Semantic Changes Indicated by Saccades

Self-supervised Temporal Discriminative Learning for Video Representation Learning

Memory-augmented Dense Predictive Coding for Video Representation Learning

Self-supervised Motion Learning from Static Images

Learning Effective Geometry Representation from Videos for Self-Supervised Monocular Depth Estimation

Self-supervised Spatiotemporal Representation Learning by Exploiting Video Continuity