Abstract:Most existing Convolutional Neural Networks(CNNs) used for action recognition are either difficult to optimize or underuse crucial temporal information. Inspired by the fact that the recurrent model consistently makes breakthroughs in the task related to sequence, we propose a novel Multi-Level Recurrent Residual Networks(MRRN) which incorporates three recognition streams. Each stream consists of a Residual Networks(ResNets) and a recurrent model. The proposed model captures spatiotemporal information by employing both alternative ResNets to learn spatial representations from static frames and stacked Simple Recurrent Units(SRUs) to model temporal dynamics. Three distinct-level streams learned low-, mid-, high-level representations independently are fused by computing a weighted average of their softmax scores to obtain the complementary representations of the video. Unlike previous models which boost performance at the cost of time complexity and space complexity, our models have a lower complexity by employing shortcut connection and are trained end-to-end with greater efficiency. MRRN displays significant performance improvements compared to CNN-RNN framework baselines and obtains comparable performance with the state-of-the-art, achieving 51.3% on HMDB-51 dataset and 81.9% on UCF-101 dataset although no additional data.

What problem does this paper attempt to address?

The problem that this paper attempts to solve is that in action recognition tasks, existing convolutional neural networks (CNNs) are either difficult to optimize or fail to fully utilize crucial temporal information. To overcome these challenges, the authors propose a new multi - level recurrent residual network (MRRN), aiming to capture spatio - temporal information by integrating three recognition streams, each consisting of a residual network (ResNets) and a recurrent model. The MRRN model learns spatial representations from static frames using alternating ResNets and models temporal dynamics by stacking simple recurrent units (SRUs), thereby independently learning low - level, mid - level and high - level representations at different levels, and fusing these representations by calculating the weighted average of their softmax scores to obtain complementary representations of the video. Compared with previous models, MRRN not only has improved performance, but also has lower temporal and spatial complexity and higher training efficiency by adopting shortcut connections. Specifically, the main contributions of the paper include: 1. **Analyzed the impact of different hyper - parameter settings on performance**: Through qualitative analysis, the general trend of performance was demonstrated, and a method of using identity shortcut connections in the proposed model to reduce spatial and temporal complexity was proposed. 2. **Experimentally verified the contribution of different - level features to action recognition**: Through experiments, how different - level features contribute to action recognition was shown, and the effects of different temporal pooling methods were explored. 3. **Proposed an architecture consisting of three independent sub - models**: These three sub - models are respectively called low - level, mid - level and high - level recurrent residual networks (RRN), which are used to simultaneously generate video representations at different levels and finally make predictions. 4. **Conducted a large number of experiments on two standard video action benchmark datasets (HMDB - 51 and UCF - 101)**: The experimental results show that MRRN significantly outperforms the baseline models based on the CNN - RNN framework in performance and is comparable to the state - of - the - art methods. Through these methods and experiments, the paper effectively solves the problems existing in existing action recognition models and provides a more efficient and better - performing solution.

Multi-Level Recurrent Residual Networks for Action Recognition

Learning SpatioTemporal and Motion Features in a Unified 2D Network for Action Recognition

Multi-scale residual network model combined with Global Average Pooling for action recognition

Action Recognition with Joint Attention on Multi-Level Deep Features

Human Action Recognition Based on Three-Stream Network with Frame Sequence Features

Hierarchical Multi-scale Attention Networks for Action Recognition

Temporal Distinct Representation Learning for Action Recognition

Spatiotemporal Residual Networks for Video Action Recognition

Action Recognition By Learning Deep Multi-Granular Spatio-Temporal Video Representation

An Attentional Spatial Temporal Graph Convolutional Network with Co-Occurrence Feature Learning for Action Recognition

Weighted Multi-Region Convolutional Neural Network for Action Recognition with Low-Latency Online Prediction

Human Action Recognition Based on Improved Fusion Attention CNN and RNN

Recognition of Visually Perceived Compositional Human Actions by Multiple Spatio-Temporal Scales Recurrent Neural Networks

Temporal Pyramid Pooling-Based Convolutional Neural Network for Action Recognition

Multi-modality Fusion Network for Action Recognition.

Low-Latency Human Action Recognition with Weighted Multi-Region Convolutional Neural Network

MSF-Net: A Multilevel Spatiotemporal Feature Fusion Network Combines Attention for Action Recognition

End-to-end Video-level Representation Learning for Action Recognition

Joint Network based Attention for Action Recognition

Convolutional Neural Network-Based Video Super-Resolution for Action Recognition

TS-LSTM and Temporal-Inception: Exploiting Spatiotemporal Dynamics for Activity Recognition