Abstract:Recent studies have shown that video-level representation learning is crucial to the capture and understanding of the long-range temporal structure for video action recognition. Most existing 3D convolutional neural network (CNN)-based methods for video-level representation learning are clip-based and focus only on short-term motion and appearances. These CNN-based methods lack the capacity to incorporate and model the long-range spatiotemporal representation of the underlying video and ignore the long-range video-level context during training. In this study, we propose a factorized 4D CNN architecture with attention (F4D) that is capable of learning more effective, finer-grained, long-term spatiotemporal video representations. We demonstrate that the proposed F4D architecture yields significant performance improvements over the conventional 2D, and 3D CNN architectures proposed in the literature. Experiment evaluation on five action recognition benchmark datasets, i.e., Something-Something-v1, SomethingSomething-v2, Kinetics-400, UCF101, and HMDB51 demonstrate the effectiveness of the proposed F4D network architecture for video-level action recognition.

What problem does this paper attempt to address?

The problem that this paper attempts to solve is how to effectively capture and understand long - range temporal structures in video action recognition. Existing methods based on 3D Convolutional Neural Networks (CNNs) mainly focus on short - term action and appearance features, lack the effective ability to model long - range spatio - temporal relationships in videos, and ignore long - range video - level context during the training process. This leads to poor performance when dealing with complex action recognition tasks that require long - time dependencies. To solve this problem, the paper proposes a new architecture - Factorized 4D Convolutional Neural Network (F4D), which can learn more effective and refined long - range spatio - temporal video representations. F4D further enhances the model's ability to focus on Regions of Interest (ROI) in videos by introducing an attention mechanism, thereby improving the model's representational ability and final recognition accuracy. Specifically, the main contributions of the paper include: - Proposing a factorized 4D convolution operation that can capture more complex long - range temporal dependencies and cross - segment interactions with lower training and testing errors. - Designing two attention mechanism modules, namely the Temporal Attention (TA) module and the Spatio - Temporal Attention (STA) module. These modules can guide the network to focus on the ROI in the video with almost no increase in computational cost, thereby enhancing the representation effect. - Constructing an effective and simple network architecture F4D. This architecture is achieved by inserting F4D residual blocks into the standard ResNet architecture and is easy to integrate into existing 3D CNN frameworks. Experimental results show that the F4D architecture performs well on multiple video action recognition benchmark datasets, including Something - Something - v1 and v2, Kinetics - 400, UCF101 and HMDB51, proving its effectiveness and superiority in video - level action recognition tasks.

F4D: Factorized 4D Convolutional Neural Network for Efficient Video-level Representation Learning

V4D:4D Convolutional Neural Networks for Video-level Representation Learning

Human Action Recognition using Factorized Spatio-Temporal Convolutional Networks

A Real-Time Action Representation With Temporal Encoding and Deep Compression

Short-Term Action Recognition by 3D Convolutional Neural Network with Pixel-Wise Evidences

DC3D: A Video Action Recognition Network Based on Dense Connection

D3D: Dual 3-D Convolutional Network for Real-Time Action Recognition

Dynamic Spatio-Temporal Feature Learning via Graph Convolution in 3D Convolutional Networks

Temporal Distinct Representation Learning for Action Recognition

Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition

3D-TDC: A 3D temporal dilation convolution framework for video action recognition

TEINet: Towards an Efficient Architecture for Video Recognition.

Gate-Shift-Fuse for Video Action Recognition

Enhanced Action Recognition With Visual Attribute-Augmented 3D Convolutional Neural Network

Visual Attribute-augmented Three-dimensional Convolutional Neural Network for Enhanced Human Action Recognition.

MULTI-DIRECTIONAL CONVOLUTION NETWORKS WITH SPATIAL-TEMPORAL FEATURE PYRAMID MODULE FOR ACTION RECOGNITION

F2D-SIFPNet: a Frequency 2D Slow-I-Fast-P Network for Faster Compressed Video Action Recognition

End-to-end Video-level Representation Learning for Action Recognition

2D or Not 2D? Adaptive 3D Convolution Selection for Efficient Video Recognition

Action Recognition By Learning Deep Multi-Granular Spatio-Temporal Video Representation

Learning Hierarchical Video Representation for Action Recognition