Three-Stream Action Tubelet Detector for Spatiotemporal Action Detection in Videos.

Yutang Wu,Hanli Wang,Qinyu Li
DOI: https://doi.org/10.1007/978-3-030-00767-6_28
2018-01-01
Abstract:In recent years, human action detection in videos has gained wide attention. Instead of detection frame by frame, a model named action tubelet (ACT) detector detects human actions sequence by sequence and achieves remarkable performances on both accuracy and speed in the form of two streams. In this work, a three-stream action tubelet detector (three-stream ACT detector) is proposed which adds an extra pose stream to obtain more information about human actions and fuses three streams by weighted average compared to the two-stream architecture. The experimental results on the benchmark UCF-Sports, J-HMDB and UCF-101 datasets demonstrate that the proposed threestream ACT detector framework is able to boost the performance of human action detection.
What problem does this paper attempt to address?