Multi-Level Recurrent Residual Networks for Action Recognition

Zhenxing Zheng,Gaoyun An,Qiuqi Ruan
DOI: https://doi.org/10.48550/arXiv.1711.08238
2018-01-03
Abstract:Most existing Convolutional Neural Networks(CNNs) used for action recognition are either difficult to optimize or underuse crucial temporal information. Inspired by the fact that the recurrent model consistently makes breakthroughs in the task related to sequence, we propose a novel Multi-Level Recurrent Residual Networks(MRRN) which incorporates three recognition streams. Each stream consists of a Residual Networks(ResNets) and a recurrent model. The proposed model captures spatiotemporal information by employing both alternative ResNets to learn spatial representations from static frames and stacked Simple Recurrent Units(SRUs) to model temporal dynamics. Three distinct-level streams learned low-, mid-, high-level representations independently are fused by computing a weighted average of their softmax scores to obtain the complementary representations of the video. Unlike previous models which boost performance at the cost of time complexity and space complexity, our models have a lower complexity by employing shortcut connection and are trained end-to-end with greater efficiency. MRRN displays significant performance improvements compared to CNN-RNN framework baselines and obtains comparable performance with the state-of-the-art, achieving 51.3% on HMDB-51 dataset and 81.9% on UCF-101 dataset although no additional data.
Computer Vision and Pattern Recognition
What problem does this paper attempt to address?
The problem that this paper attempts to solve is that in action recognition tasks, existing convolutional neural networks (CNNs) are either difficult to optimize or fail to fully utilize crucial temporal information. To overcome these challenges, the authors propose a new multi - level recurrent residual network (MRRN), aiming to capture spatio - temporal information by integrating three recognition streams, each consisting of a residual network (ResNets) and a recurrent model. The MRRN model learns spatial representations from static frames using alternating ResNets and models temporal dynamics by stacking simple recurrent units (SRUs), thereby independently learning low - level, mid - level and high - level representations at different levels, and fusing these representations by calculating the weighted average of their softmax scores to obtain complementary representations of the video. Compared with previous models, MRRN not only has improved performance, but also has lower temporal and spatial complexity and higher training efficiency by adopting shortcut connections. Specifically, the main contributions of the paper include: 1. **Analyzed the impact of different hyper - parameter settings on performance**: Through qualitative analysis, the general trend of performance was demonstrated, and a method of using identity shortcut connections in the proposed model to reduce spatial and temporal complexity was proposed. 2. **Experimentally verified the contribution of different - level features to action recognition**: Through experiments, how different - level features contribute to action recognition was shown, and the effects of different temporal pooling methods were explored. 3. **Proposed an architecture consisting of three independent sub - models**: These three sub - models are respectively called low - level, mid - level and high - level recurrent residual networks (RRN), which are used to simultaneously generate video representations at different levels and finally make predictions. 4. **Conducted a large number of experiments on two standard video action benchmark datasets (HMDB - 51 and UCF - 101)**: The experimental results show that MRRN significantly outperforms the baseline models based on the CNN - RNN framework in performance and is comparable to the state - of - the - art methods. Through these methods and experiments, the paper effectively solves the problems existing in existing action recognition models and provides a more efficient and better - performing solution.