MuseTalk: Real-Time High Quality Lip Synchronization with Latent Space Inpainting

Yue Zhang,Minhao Liu,Zhaokang Chen,Bin Wu,Yubin Zeng,Chao Zhan,Yingjie He,Junxin Huang,Wenjiang Zhou

2024-10-16

Abstract:Achieving high-resolution, identity consistency, and accurate lip-speech synchronization in face visual dubbing presents significant challenges, particularly for real-time applications like live video streaming. We propose MuseTalk, which generates lip-sync targets in a latent space encoded by a Variational Autoencoder, enabling high-fidelity talking face video generation with efficient inference. Specifically, we project the occluded lower half of the face image and itself as an reference into a low-dimensional latent space and use a multi-scale U-Net to fuse audio and visual features at various levels. We further propose a novel sampling strategy during training, which selects reference images with head poses closely matching the target, allowing the model to focus on precise lip movement by filtering out redundant information. Additionally, we analyze the mechanism of lip-sync loss and reveal its relationship with input information volume. Extensive experiments show that MuseTalk consistently outperforms recent state-of-the-art methods in visual fidelity and achieves comparable lip-sync accuracy. As MuseTalk supports the online generation of face at 256x256 at more than 30 FPS with negligible starting latency, it paves the way for real-time applications.

Computer Vision and Pattern Recognition

What problem does this paper attempt to address?

The paper attempts to address the problem of generating high-resolution, identity-consistent, and lip-synced facial visual dubbing videos in real-time applications. Specifically, the paper focuses on the following three main challenges: 1. **High Resolution**: The generated facial video needs to maintain high resolution to ensure the authenticity and detail of the visual effects. 2. **Identity Consistency**: The generated facial video needs to maintain identity consistency with the original face, avoiding distortion of facial features due to errors in the generation process. 3. **Lip Syncing**: The generated facial video needs to be highly synchronized with the input audio, ensuring that lip movements precisely match the speech content. These challenges are particularly significant in real-time applications, such as live video streaming. Existing methods have some shortcomings in addressing these issues, such as low generation quality, high computational resource consumption, and slow inference speed. Therefore, this paper proposes a new method called MuseTalk, which aims to solve the above problems by performing lip inpainting in the latent space, achieving efficient and high-quality real-time facial visual dubbing.

MuseTalk: Real-Time High Quality Lip Synchronization with Latent Space Inpainting

SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory

High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion Model

Audio-driven Talking Face Video Generation with Natural Head Pose

VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the Wild

Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short Video

SyncTalklip: Highly Synchronized Lip-Readable Speaker Generation with Multi-Task Learning

Real-time Lip Synchronization Based on Hidden Markov Models

Meta Talk: Learning To Data-Efficiently Generate Audio-Driven Lip-Synchronized Talking Face With High Definition

VividWav2Lip: High-Fidelity Facial Animation Generation Based on Speech-Driven Lip Synchronization

Make Your Actor Talk: Generalizable and High-Fidelity Lip Sync with Motion and Appearance Disentanglement

LaDTalk: Latent Denoising for Synthesizing Talking Head Videos with High Frequency Details

Style-Preserving Lip Sync via Audio-Aware Style Reference

MILG: Realistic Lip-Sync Video Generation with Audio-Modulated Image Inpainting

Multimodal Learning for Temporally Coherent Talking Face Generation with Articulator Synergy

RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network

Towards Realistic Visual Dubbing with Heterogeneous Sources

Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts

VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Real-Time Lip Sync for Live 2D Animation

PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head Generation