Abstract:Parametric and non-parametric methods are two commonly used strategies in current 3D hand pose reconstruction. Parametric methods predict low-dimensional parameters to fit a predefined hand model to the input image. Benefiting from the prior knowledge of hand models, parametric methods guarantee plausible hand poses, whereas the pose estimation accuracy is limited due to nonlinear regression and spatial information loss. Differently, non-parametric methods directly estimate the coordinates of keypoints or mesh vertices from the input image. The reconstructed 3D hand poses show high precision but may be less robust. In this paper, we integrate the advantages of two methods for accurate and robust hand pose reconstruction. Specifically, we disentangle the hand pose reconstruction into global modeling and local refinement, which is performed in a coarse-to-fine manner. Firstly, we utilize global features from the encoder to generate the initial estimation by a parametric method, which aims to provide the prior knowledge of the human hand for subsequent processes. Then, we gradually fuse multi-scale contextual features for local refinement by explicitly integrating global prior information and local visual features. In particular, we introduce a consecutive pixel-aligned feature retrieval module to extract fine-grained information from visual features, thereby achieving pixel-level alignment. Furthermore, we demonstrate that our method can be extended to weakly-supervised learning where only sparse pose annotations are needed, potentially alleviating the burden of meticulous mesh annotation. The effectiveness and robustness of our method are substantiated through both fully- and weakly-supervised experiments, demonstrating superior performance compared to state-of-the-art methods. We plan to release our code at https://github.com/Kun-Gao/P_GLFnet.

Decoupling Heterogeneous Features for Robust 3D Interacting Hand Poses Estimation

FDN: Feature Decoupling Network for Head Pose Estimation.

Decoupled Iterative Refinement Framework for Interacting Hands Reconstruction from a Single RGB Image

CAMInterHand: Cooperative Attention for Multi-View Interactive Hand Pose and Mesh Reconstruction

Dual Regression for Efficient Hand Pose Estimation

3D Interacting Hand Pose Estimation by Hand De-occlusion and Removal

A hybrid network for estimating 3D interacting hand pose from a single RGB image

Denoising Diffusion for 3D Hand Pose Estimation from Images

Disentangling Pose from Appearance in Monochrome Hand Images

Joint Hand-Object Pose Estimation with Differentiably-Learned Physical Contact Point Analysis

Joint-wise 2D to 3D lifting for hand pose estimation from a single RGB image

Weakly Supervised Segmentation Guided Hand Pose Estimation During Interaction with Unknown Objects.

3D Hand Reconstruction via Aggregating Intra and Inter Graphs Guided by Prior Knowledge for Hand-Object Interaction Scenario

LHFF-Net: Local heterogeneous feature fusion network for 6DoF pose estimation

Progressively Global-Local Fusion with Explicit Guidance for Accurate and Robust 3d Hand Pose Reconstruction

Which patients should undergo laparoscopy?

Recurrent 3D Hand Pose Estimation Using Cascaded Pose-Guided 3D Alignments

Robust 3D Hand Detection from a Single RGB-D Image in Unconstrained Environments

A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB Image

HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields