Abstract:Multi-frame human pose estimation has long been an appealing and fundamental issue in visual perception. Owing to the frequent rapid motion and pose occlusion in videos, this task is extremely challenging. Current state-of-the-art methods seek to model spatiotemporal features by equally fusing each frame in the local sequence, which weakens the target frame information. In addition, existing approaches usually emphasize more on deep features while ignoring the detailed information implied in the shallow feature maps, resulting in the dropping of crucial features. To address the above problems, we propose an effective framework, namely spatiotemporal learning transformer for video-based human pose estimation (SLT-Pose), which consists of a Personalized Feature Extraction Module (PFEM), Self-feature Refinement Module (SRM), Cross-frame Temporal Learning Module (CTLM) and Disentangled Keypoint Detector (DKD). To be specific, we propose PFEM which extracts and modulates the individual frame features to adapt to the varying human shape, and integrates single-frame features to obtain the spatiotemporal features. We further present SRM to establish global correlation spatial cues on the target frame to attain the refinement feature. Then, a CTLM is designed to search for the information most closely related to the target frame from the spatiotemporal features to intensify the interaction between the target frame and the local sequence, using both the shallow detailed and the deep semantic representations. Finally, we employ DKD to extract the disentangled characteristics of each joint and encode the articulated joint pairs in the human body, promoting the model to reasonably and accurately predict the keypoint heatmaps. Extensive experiments on three huamn motion benchmarks, including PoseTrack2017, PoseTrack2018, and Sub-JHMDB dataset, demonstrate that SLT-Pose plays favorably against state-of-the-art approaches in terms of both objective evaluation and subjective visual performance.

PGVT: Pose-Guided Video Transformer for Fine-Grained Action Recognition

Spatiotemporal Learning Transformer for Video-Based Human Pose Estimation

Seeing the Pose in the Pixels: Learning Pose-Aware Representations in Vision Transformers

PSVT: End-to-End Multi-person 3D Pose and Shape Estimation with Progressive Video Transformers.

IVT: An End-to-End Instance-guided Video Transformer for 3D Pose Estimation

TP-VIT: A Two-Pathway Vision Transformer for Video Action Recognition

Joint Multi-Scale Transformers and Pose Equivalence Constraints for 3D Human Pose Estimation

A human activity recognition method based on Vision Transformer

PeVL: Pose-Enhanced Vision-Language Model for Fine-Grained Human Action Recognition

Pyramid Spatial-Temporal Graph Transformer for Skeleton-Based Action Recognition

ViTPose++: Vision Transformer for Generic Body Pose Estimation

Bilateral Pose Transformer for Human Pose Estimation.

GITPose: going shallow and deeper using vision transformers for human pose estimation

MM-ViT: Multi-Modal Video Transformer for Compressed Video Action Recognition

MgMViT: Multi-Granularity and Multi-Scale Vision Transformer for Efficient Action Recognition

Decoupled Spatio-Temporal Grouping Transformer for Skeleton-Based Action Recognition

HRPVT: High-Resolution Pyramid Vision Transformer for medium and small-scale human pose estimation

Relative-position Embedding Based Spatially and Temporally Decoupled Transformer for Action Recognition

HEViTPose: High-Efficiency Vision Transformer for Human Pose Estimation

Multiple View Geometry Transformers for 3D Human Pose Estimation

LS-VIT: Vision Transformer for action recognition based on long and short-term temporal difference