Multi-cue based four-stream 3D ResNets for video-based action recognition

Lei Wang,Xiaoguang Yuan,Ming Zong,Yujun Ma,Wanting Ji,Mingzhe Liu,Ruili Wang
DOI: https://doi.org/10.1016/j.ins.2021.07.079
IF: 8.1
2021-10-01
Information Sciences
Abstract:Action recognition is one of the important computer vision tasks, which has many applications. This paper proposes a Multi-cue based Four-stream 3D ResNets (MF3D) model for action recognition. The proposed MF3D model contains four streams: a video saliency stream, an appearance stream, a motion stream and an audio stream. Four cues (i.e. the appearance cue, the motion cue, the video saliency cue and audio cue) are captured by the four streams of our proposed MF3D model. In addition, three different connections between different streams are injected, which can transfer different cues between different streams to obtain more effective spatiotemporal features. Experiments are conducted on the Kinetics and Kinetics-Sounds datasets, and the results verify that our MF3D model is effective and outperforms current existing models.
computer science, information systems
What problem does this paper attempt to address?