Abstract:The generation of co-speech gestures for digital humans is an emerging area in the field of virtual human creation. Prior research has progressed by using acoustic and semantic information as input and adopting a classification method to identify the person's ID and emotion for driving co-speech gesture generation. However, this endeavor still faces significant challenges. These challenges go beyond the intricate interplay among co-speech gestures, speech acoustic, and semantics; they also encompass the complexities associated with personality, emotion, and other obscure but important factors. This paper introduces "DiT-Gestures", a speech-conditional diffusion-based and non-autoregressive transformer-based generative model with the WavLM pre-trained model and a dynamic mask attention network (DMAN). It can produce individual and stylized full-body co-speech gestures by only using raw speech audio, eliminating the need for complex multimodal processing and manual annotation. Firstly, considering that speech audio contains acoustic and semantic features and conveys personality traits, emotions, and more subtle information related to accompanying gestures, we pioneer the adaptation of WavLM, a large-scale pre-trained model, to extract the style from raw audio information. Secondly, we replace the causal mask by introducing a learnable dynamic mask for better local modeling in the neighborhood of the target frames. Extensive subjective evaluation experiments are conducted on the Trinity, ZEGGS, and BEAT datasets to confirm WavLM's and the model's ability to synthesize natural co-speech gestures with various styles.

DiffuseStyleGesture: Stylized Audio-Driven Co-Speech Gesture Generation with Diffusion Models

Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation

Conversational Co-Speech Gesture Generation via Modeling Dialog Intention, Emotion, and Context with Diffusion Models

The DiffuseStyleGesture+ entry to the GENEA Challenge 2023

DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation

Freetalker: Controllable Speech and Text-Driven Gesture Generation Based on Diffusion Models for Enhanced Speaker Naturalness

SIGGesture: Generalized Co-Speech Gesture Synthesis via Semantic Injection with Large-Scale Pre-Training Diffusion Models

Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model

Multimodal-driven Talking Face Generation, Face Swapping, Diffusion Model

ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance

DiffMotion: Speech-Driven Gesture Synthesis Using Denoising Diffusion Model

C2G2: Controllable Co-speech Gesture Generation with Latent Diffusion Model

GestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents

DiT-Gesture: A Speech-Only Approach to Stylized Gesture Generation

DiffPoseTalk: Speech-Driven Stylistic 3D Facial Animation and Head Pose Generation via Diffusion Models

Cultural Self-Adaptive Multimodal Gesture Generation Based on Multiple Culture Gesture Dataset

UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons