A Two-stream Hybrid CNN-Transformer Network for Skeleton-based Human Interaction Recognition

Ruoqi Yin,Jianqin Yin
2023-12-31
Abstract:Human Interaction Recognition is the process of identifying interactive actions between multiple participants in a specific situation. The aim is to recognise the action interactions between multiple entities and their meaning. Many single Convolutional Neural Network has issues, such as the inability to capture global instance interaction features or difficulty in training, leading to ambiguity in action semantics. In addition, the computational complexity of the Transformer cannot be ignored, and its ability to capture local information and motion features in the image is poor. In this work, we propose a Two-stream Hybrid CNN-Transformer Network (THCT-Net), which exploits the local specificity of CNN and models global dependencies through the Transformer. CNN and Transformer simultaneously model the entity, time and space relationships between interactive entities respectively. Specifically, Transformer-based stream integrates 3D convolutions with multi-head self-attention to learn inter-token correlations; We propose a new multi-branch CNN framework for CNN-based streams that automatically learns joint spatio-temporal features from skeleton sequences. The convolutional layer independently learns the local features of each joint neighborhood and aggregates the features of all joints. And the raw skeleton coordinates as well as their temporal difference are integrated with a dual-branch paradigm to fuse the motion features of the skeleton. Besides, a residual structure is added to speed up training convergence. Finally, the recognition results of the two branches are fused using parallel splicing. Experimental results on diverse and challenging datasets, demonstrate that the proposed method can better comprehend and infer the meaning and context of various actions, outperforming state-of-the-art methods.
Computer Vision and Pattern Recognition,Artificial Intelligence
What problem does this paper attempt to address?
The problem that this paper attempts to solve is the challenges in Human Interaction Recognition (HIR). Specifically, the paper focuses on how to accurately identify and understand the interactive behaviors among multiple people as well as their meanings and purposes in video sequences. The existing single Convolutional Neural Network (CNN) has problems in capturing global instance interaction features or training difficulty, resulting in ambiguity in action semantics. Although the Transformer can capture long - distance dependencies, its computational complexity is high, and its ability to capture local information in images and motion features is weak. To solve these problems, the author proposes a Two - stream Hybrid CNN - Transformer Network (THCT - Net). This network combines the CNN's ability to capture local characteristics and the Transformer's ability to model global dependencies, and simultaneously processes the relationships of entities, time and space. Specifically: 1. **Transformer branch**: Learn the correlations between tokens by integrating 3D convolution and multi - head self - attention mechanism. 2. **CNN branch**: Propose a new multi - branch CNN framework to automatically learn joint spatio - temporal features from skeleton sequences. The convolutional layer independently learns the local features of each joint neighborhood and aggregates the features of all joints. The original skeleton coordinates and their time differences are fused through a two - branch paradigm to fuse the motion features of the skeleton. In addition, a residual structure is added to accelerate training convergence. Finally, the recognition results of the two branches are fused by parallel concatenation. Experimental results show that this method can better understand and infer the meanings and contexts of various actions on multiple challenging datasets (such as NTU - RGBD, H2O and Assembly101), and is superior to the existing state - of - the - art methods.