Two-Stream Convolutional Neural Network for Video Action Recognition
Han Qiao,Shuang Liu,Qingzhen Xu,Shouqiang Liu,Wanggan Yang
DOI: https://doi.org/10.3837/tiis.2021.10.011
2021-01-01
KSII Transactions on Internet and Information Systems
Abstract:Video action recognition is widely used in video surveillance, behavior detection, human-computer interaction, medically assisted diagnosis and motion analysis. However, video action recognition can be disturbed by many factors, such as background, illumination and so on. Two-stream convolutional neural network uses the video spatial and temporal models to train separately, and performs fusion at the output end. The multi segment Two-Stream convolutional neural network model trains temporal and spatial information from the video to extract their feature and fuse them, then determine the category of video action. Google Xception model and the transfer learning is adopted in this paper, and the Xception model which trained on ImageNet is used as the initial weight. It greatly overcomes the problem of model underfitting caused by insufficient video behavior dataset , and it can effectively reduce the influence of various factors in the video. This way also greatly improves the accuracy and reduces the training time. What’s more, to make up for the shortage of dataset, the kinetics400 dataset was used for pre-training, which greatly improved the accuracy of the model. In this applied research, through continuous efforts, the expected goal is basically achieved, and according to the study and research, the design of the original dual-flow model is improved.