Abstract:Skeleton-based human interaction recognition is a challenging task in the field of vision and image processing. Graph Convolutional Networks (GCNs) achieved remarkable performance by modeling the human skeleton as a topology. However, existing GCN-based methods have two problems: (1) Existing frameworks cannot effectively take advantage of the complementary features of different skeletal modalities. There is no information transfer channel between various specific modalities. (2) Limited by the structure of the skeleton topology, it is hard to capture and learn the information about two-person interactions. To solve these problems, inspired by the human visual neural network, we propose a multi-modal enhancement transformer (ME-Former) network for skeleton-based human interaction recognition. ME-Former includes a multi-modal enhancement module (ME) and a context progressive fusion block (CPF). More specifically, each ME module consists of a multi-head cross-modal attention block (MH-CA) and a two-person hypergraph self-attention block (TH-SA), which are responsible for enhancing the skeleton features of a specific modality from other skeletal modalities and modeling spatial dependencies between joints using the specific modality, respectively. In addition, we propose a two-person skeleton topology and a two-person hypergraph representation. The TH-SA block can embed their structural information into the self-attention to better learn two-person interaction. The CPF block is capable of progressively transforming the features of different skeletal modalities from low-level features to higher-order global contexts, making the enhancement process more efficient. Extensive experiments on benchmark NTU-RGB+D 60 and NTU-RGB+D 120 datasets consistently verify the effectiveness of our proposed ME-Former by outperforming state-of-the-art methods.

MLDT: Multi-task Learning with Denoising Transformer for Gait Identity and Emotion Recognition

A Multi-Head Pseudo Nodes Based Spatial–temporal Graph Convolutional Network for Emotion Perception from GAIT

Exploring Self-Supervised Vision Transformers for Gait Recognition in the Wild

Multi-scale Context-aware Network with Transformer for Gait Recognition

Multi-Modal Transformer with Skeleton and Text for Action Recognition

Spatial Transformer Network on Skeleton‐based Gait Recognition

Gait-CNN-ViT: Multi-Model Gait Recognition with Convolutional Neural Networks and Vision Transformer

Disentangling 3D/4D Facial Affect Recognition with Faster Multi-View Transformer

Multi-Scale Adaptive Skeleton Transformer for action recognition

TNTC: Two-Stream Network with Transformer-Based Complementarity for Gait-Based Emotion Recognition

GaitGMT: Global feature mapping transformer for gait recognition

Exploring Transformers for Behavioural Biometrics: A Case Study in Gait Recognition

Multi-Modal Enhancement Transformer Network for Skeleton-Based Human Interaction Recognition

DyGait: Exploiting Dynamic Representations for High-performance Gait Recognition

STDM-transformer: Space-time dual multi-scale transformer network for skeleton-based action recognition

A human activity recognition method based on Vision Transformer

A Novel Two-Stream Transformer-Based Framework for Multi-Modality Human Action Recognition

MVTN: A Multiscale Video Transformer Network for Hand Gesture Recognition

3Mformer: Multi-order Multi-mode Transformer for Skeletal Action Recognition

HorGait: Advancing Gait Recognition with Efficient High-Order Spatial Interactions in LiDAR Point Clouds

On Learning Disentangled Representations for Gait Recognition