Abstract:Human Interaction Recognition is the process of identifying interactive actions between multiple participants in a specific situation. The aim is to recognise the action interactions between multiple entities and their meaning. Many single Convolutional Neural Network has issues, such as the inability to capture global instance interaction features or difficulty in training, leading to ambiguity in action semantics. In addition, the computational complexity of the Transformer cannot be ignored, and its ability to capture local information and motion features in the image is poor. In this work, we propose a Two-stream Hybrid CNN-Transformer Network (THCT-Net), which exploits the local specificity of CNN and models global dependencies through the Transformer. CNN and Transformer simultaneously model the entity, time and space relationships between interactive entities respectively. Specifically, Transformer-based stream integrates 3D convolutions with multi-head self-attention to learn inter-token correlations; We propose a new multi-branch CNN framework for CNN-based streams that automatically learns joint spatio-temporal features from skeleton sequences. The convolutional layer independently learns the local features of each joint neighborhood and aggregates the features of all joints. And the raw skeleton coordinates as well as their temporal difference are integrated with a dual-branch paradigm to fuse the motion features of the skeleton. Besides, a residual structure is added to speed up training convergence. Finally, the recognition results of the two branches are fused using parallel splicing. Experimental results on diverse and challenging datasets, demonstrate that the proposed method can better comprehend and infer the meaning and context of various actions, outperforming state-of-the-art methods.

What problem does this paper attempt to address?

The problem that this paper attempts to solve is the challenges in Human Interaction Recognition (HIR). Specifically, the paper focuses on how to accurately identify and understand the interactive behaviors among multiple people as well as their meanings and purposes in video sequences. The existing single Convolutional Neural Network (CNN) has problems in capturing global instance interaction features or training difficulty, resulting in ambiguity in action semantics. Although the Transformer can capture long - distance dependencies, its computational complexity is high, and its ability to capture local information in images and motion features is weak. To solve these problems, the author proposes a Two - stream Hybrid CNN - Transformer Network (THCT - Net). This network combines the CNN's ability to capture local characteristics and the Transformer's ability to model global dependencies, and simultaneously processes the relationships of entities, time and space. Specifically: 1. **Transformer branch**: Learn the correlations between tokens by integrating 3D convolution and multi - head self - attention mechanism. 2. **CNN branch**: Propose a new multi - branch CNN framework to automatically learn joint spatio - temporal features from skeleton sequences. The convolutional layer independently learns the local features of each joint neighborhood and aggregates the features of all joints. The original skeleton coordinates and their time differences are fused through a two - branch paradigm to fuse the motion features of the skeleton. In addition, a residual structure is added to accelerate training convergence. Finally, the recognition results of the two branches are fused by parallel concatenation. Experimental results show that this method can better understand and infer the meanings and contexts of various actions on multiple challenging datasets (such as NTU - RGBD, H2O and Assembly101), and is superior to the existing state - of - the - art methods.

A Two-stream Hybrid CNN-Transformer Network for Skeleton-based Human Interaction Recognition

Multi-Modal Enhancement Transformer Network for Skeleton-Based Human Interaction Recognition

Multi-Stream Interaction Networks for Human Action Recognition

Two-Stream 3D Convolutional Neural Network for Skeleton-Based Action Recognition

Human Behavior Recognition Based on CNN-LSTM Hybrid and Multi-Sensing Feature Information Fusion

A Skeleton-Based Assembly Action Recognition Method with Feature Fusion for Human-Robot Collaborative Assembly

HybridNet: Integrating GCN and CNN for skeleton-based action recognition

A Novel Two-Stream Transformer-Based Framework for Multi-Modality Human Action Recognition

Skeleton-Based Human Action Recognition Using Spatial Temporal 3D Convolutional Neural Networks

Two-stream Multi-level Dynamic Point Transformer for Two-person Interaction Recognition

Spatial Temporal Transformer Network for Skeleton-based Action Recognition

Human Action Recognition Based on Three-Stream Network with Frame Sequence Features

Human-Robot Collaboration Through a Multi-Scale Graph Convolution Neural Network With Temporal Attention

Interactive semantics neural networks for skeleton-based human interaction recognition

MSST-RT: Multi-Stream Spatial-Temporal Relative Transformer for Skeleton-Based Action Recognition

HDBN: A Novel Hybrid Dual-branch Network for Robust Skeleton-based Action Recognition

Cmf-transformer: cross-modal fusion transformer for human action recognition

Multi-Modal Transformer with Skeleton and Text for Action Recognition

Temporal Enhanced Multi-Stream Graph Convolutional Nerual Networks For Skeleton-Based Action Recognition

Symmetrical Enhanced Fusion Network for Skeleton-Based Action Recognition

High Efficient LSTM-based Network for Human Interaction Understanding