Abstract:Creating pose-driven human avatars is about modeling the mapping from the low-frequency driving pose to high-frequency dynamic human appearances, so an effective pose encoding method that can encode high-fidelity human details is essential to human avatar modeling. To this end, we present PoseVocab, a novel pose encoding method that encourages the network to discover the optimal pose embeddings for learning the dynamic human appearance. Given multi-view RGB videos of a character, PoseVocab constructs key poses and latent embeddings based on the training poses. To achieve pose generalization and temporal consistency, we sample key rotations in $so(3)$ of each joint rather than the global pose vectors, and assign a pose embedding to each sampled key rotation. These joint-structured pose embeddings not only encode the dynamic appearances under different key poses, but also factorize the global pose embedding into joint-structured ones to better learn the appearance variation related to the motion of each joint. To improve the representation ability of the pose embedding while maintaining memory efficiency, we introduce feature lines, a compact yet effective 3D representation, to model more fine-grained details of human appearances. Furthermore, given a query pose and a spatial position, a hierarchical query strategy is introduced to interpolate pose embeddings and acquire the conditional pose feature for dynamic human synthesis. Overall, PoseVocab effectively encodes the dynamic details of human appearance and enables realistic and generalized animation under novel poses. Experiments show that our method outperforms other state-of-the-art baselines both qualitatively and quantitatively in terms of synthesis quality. Code is available at <a class="link-external link-https" href="https://github.com/lizhe00/PoseVocab" rel="external noopener nofollow">this https URL</a>.

VoCAPTER: Voting-based Pose Tracking for Category-level Articulated Object Via Inter-frame Priors

CAPTRA: CAtegory-level Pose Tracking for Rigid and Articulated Objects from Point Clouds

Multi-modal 3D Human Tracking for Robots in Complex Environment with Siamese Point-Video Transformer

Temporal Consistent Object Pose Estimation from Monocular Videos

3D Point-to-Keypoint Voting Network for 6D Pose Estimation

PA-Pose: Partial Point Cloud Fusion Based on Reliable Alignment for 6D Pose Tracking

Category-Independent Articulated Object Tracking with Factor Graphs

3D Articulated Hand Tracking Based on Behavioral Model

Articulated Object Manipulation using Online Axis Estimation with SAM2-Based Tracking

Towards Real-World Category-level Articulation Pose Estimation

Corr-Track: Category-Level 6D Pose Tracking with Soft-Correspondence Matrix Estimation

Towards Real-World Aerial Vision Guidance with Categorical 6D Pose Tracker

2D-3D Pose Tracking with Multi-View Constraints

Multi-Person Articulated Tracking With Spatial and Temporal Embeddings

Learning Variational Motion Prior for Video-based Motion Capture

vmTracking: Virtual Markers Overcome Occlusion and Crowding in Multi-Animal Pose Tracking

PoseVocab: Learning Joint-structured Pose Embeddings for Human Avatar Modeling

Articulated point pattern matching in optical motion capture systems

Enhanced Multi-Object Tracking Using Pose-based Virtual Markers in 3x3 Basketball

CVAM-Pose: Conditional Variational Autoencoder for Multi-Object Monocular Pose Estimation

ContactArt: Learning 3D Interaction Priors for Category-level Articulated Object and Hand Poses Estimation