Skimming and Scanning for Efficient Action Recognition in Untrimmed Videos

Yunyan Hong,Ailing Zeng,Min Li,Cewu Lu,Li Jian,Qiang Xu
DOI: https://doi.org/10.1109/cisp-bmei53629.2021.9624415
2021-01-01
Abstract:Video action recognition (VAR) aims to classify videos into a predefined set of classes, which is a primary task of video understanding. We mainly focus on the VAR of untrimmed videos because they are most common videos in real-life scenes. Untrimmed videos have redundant and diverse clips containing contextual information, so sampling the clips is essential. Recently, some works attempt to train a generic model to select the $N$ most representative clips. However, it is difficult to model the complex relations from intra-class clips and inter-class videos within a single model and fixed selected number, and the entanglement of multiple relations is also hard to explain. Thus, instead of “only look once”, we argue “divide and conquer” strategy will be more suitable in untrimmed VAR. Inspired by the speed reading mechanism, we propose a simple yet effective clip-level solution based on skim-scan techniques. Specifically, the proposed Skim-Scan framework first skims the entire video and drops those uninformative and misleading clips. For the remaining clips, it scans clips with diverse features gradually to drop redundant clips but cover essential content. The above strategies can adaptively select the necessary clips according to the difficulty of the different videos. In order to further cut computational overhead, we observe the similar statistical expression between lightweight and heavy networks. Thus, we explore the combination of them to trade off the computational complexity and performance. Comprehensive experiments are performed on ActivityNet and mini-FCVID datasets, and results demonstrate that our solution surpasses the state-of-the-art performance in terms of accuracy and efficiency.
What problem does this paper attempt to address?