Abstract:The creation of lifelike speech-driven 3D facial animation requires a natural and precise synchronization between audio input and facial expressions. However, existing works still fail to render shapes with flexible head poses and natural facial details (e.g., wrinkles). This limitation is mainly due to two aspects: 1) Collecting training set with detailed 3D facial shapes is highly expensive. This scarcity of detailed shape annotations hinders the training of models with expressive facial animation. 2) Compared to mouth movement, the head pose is much less correlated to speech content. Consequently, concurrent modeling of both mouth movement and head pose yields the lack of facial movement controllability. To address these challenges, we introduce VividTalker, a new framework designed to facilitate speech-driven 3D facial animation characterized by flexible head pose and natural facial details. Specifically, we explicitly disentangle facial animation into head pose and mouth movement and encode them separately into discrete latent spaces. Then, these attributes are generated through an autoregressive process leveraging a window-based Transformer architecture. To augment the richness of 3D facial animation, we construct a new 3D dataset with detailed shapes and learn to synthesize facial details in line with speech content. Extensive quantitative and qualitative experiments demonstrate that VividTalker outperforms state-of-the-art methods, resulting in vivid and realistic speech-driven 3D facial animation.

Facial Animation System for Embedded Application

An Online Speech Driven Talking Head System

Speech Driven Facial Animation Using Chinese Mandarin Pronunciation Rules

Text-driven Visual Prosody Generation for Embodied Conversational Agents

Emotional Chinese talking head system

Real-time Speech-Driven Animation of Expressive Talking Faces.

Real-time Synthesis of Chinese Visua using MPEG-4 FAP Features in a

Real-time Synthesis of Chinese Visual Speech and Facial Expressions Using MPEG-4 FAP Features in a Three-Dimensional Avatar

3D Facial Animation from Chinese Text.

Animating a Chinese interactive virtual character

Individual 3D Face Synthesis Based on Orthogonal Photos and Speech-Driven Facial Animation

Creative cartoon face synthesis system for mobile entertainment

An MPEG-4 Compliant Speech Animation System

Text to Avatar in Multi-modal Human Computer Interface

Speech-driven Cartoon Animation with Emotions

Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation

3D Realistic Talking Face Co-Driven by Text and Speech

Breathing Life into Faces: Speech-driven 3D Facial Animation with Natural Head Pose and Detailed Shape

Dynamic mapping method based speech driven face animation system

Individual facial image synthesis system for a virtual human

Speech-driven Facial Animation with Spectral Gathering and Temporal Attention.