Abstract:Human motion copy is an intriguing yet challenging task in artificial intelligence and computer vision, which strives to generate a fake video of a target person performing the motion of a source person. The problem is inherently challenging due to the subtle human-body texture details to be generated and the temporal consistency to be considered. Existing approaches typically adopt a conventional GAN with an L1 or L2 loss to produce the target fake video, which intrinsically necessitates a large number of training samples that are challenging to acquire. Meanwhile, current methods still have difficulties in attaining realistic image details and temporal consistency, which unfortunately can be easily perceived by human observers. Motivated by this, we try to tackle the issues from three aspects: (1) We constrain pose-to-appearance generation with a perceptual loss and a theoretically motivated Gromov-Wasserstein loss to bridge the gap between pose and appearance. (2) We present an episodic memory module in the pose-to-appearance generation to propel continuous learning that helps the model learn from its past poor generations. We also utilize geometrical cues of the face to optimize facial details and refine each key body part with a dedicated local GAN. (3) We advocate generating the foreground in a sequence-to-sequence manner rather than a single-frame manner, explicitly enforcing temporal inconsistency. Empirical results on five datasets, iPER, ComplexMotion, SoloDance, Fish, and Mouse datasets, demonstrate that our method is capable of generating realistic target videos while precisely copying motion from a source video. Our method significantly outperforms state-of-the-art approaches and gains 7.2% and 12.4% improvements in PSNR and FID respectively.

Structure-Constrained Motion Sequence Generation.

MotionVideoGAN: A Novel Video Generator Based on the Motion Space Learned from Image Pairs

High-Quality Video Generation from Static Structural Annotations.

Human Motion Generation Via Cross-Space Constrained Sampling.

Structure-Aware Human-Action Generation

Multi-person/Group Interactive Video Generation

Multi-object Video Generation from Single Frame Layouts.

Motion Prompting: Controlling Video Generation with Motion Trajectories

Skeleton-Aided Articulated Motion Generation

Controllable GAN Synthesis Using Non-Rigid Structure-from-Motion

Multi-person/Group Interactive Video Generation

Pose Guided Human Video Generation

Deep Video Generation, Prediction and Completion of Human Action Sequences

Action2video: Generating Videos of Human 3D Actions

Action-conditioned video data improves predictability

Motion Dreamer: Realizing Physically Coherent Video Generation through Scene-Aware Motion Reasoning

Structure Preserving Video Prediction

ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation

Generative Tweening: Long-term Inbetweening of 3D Human Motions

Towards Smooth Video Composition

Do as I Do: Pose Guided Human Motion Copy