Abstract:While considerable progress has been made in achieving accurate lip synchronization for 3D speech-driven talking face generation, the task of incorporating expressive facial detail synthesis aligned with the speaker's speaking status remains challenging. Existing efforts either focus on learning a dynamic talking head pose synchronized with speech rhythm or aim for stylized facial movements guided by external reference such as emotional labels or reference video clips. The former works often yield coarse alignment, neglecting the emotional nuances present in the audio content while the latter studies lead to unnatural applications, requiring manual style source selection by users. Our goal is to directly leverage the inherent style information conveyed by human speech for generating an expressive talking face that aligns with the speaking status. In this paper, we propose AVI-Talking, an Audio-Visual Instruction system for expressive Talking face generation. This system harnesses the robust contextual reasoning and hallucination capability offered by Large Language Models (LLMs) to instruct the realistic synthesis of 3D talking faces. Instead of directly learning facial movements from human speech, our two-stage strategy involves the LLMs first comprehending audio information and generating instructions implying expressive facial details seamlessly corresponding to the speech. Subsequently, a diffusion-based generative network executes these instructions. This two-stage process, coupled with the incorporation of LLMs, enhances model interpretability and provides users with flexibility to comprehend instructions and specify desired operations or modifications. Specifically, given a speech clip, we first employ a Q-Former for contrastive alignment the speech features with visual instructions, which is then projected to input text embedding of LLMs. It functions as a prompting strategy, prompting LLMs to generate plausible visual instructions that encompass diverse facial details. In order to use these predicted instructions, a language-guided talking face generation system with disentangled latent space is delicately derived, where the speech content related lip movements and emotion correlated facial expressions are separately represented in speech content space and content irrelevant space. Additionally, we introduce a contrastive instruction-style alignment and diffusion technique within the content-irrelevant space to fully exploit the talking prior network for diverse instruction-following synthesis. Extensive experiments showcase the effectiveness of our approach in producing vivid talking faces with expressive facial movements and consistent emotional status.

Low Level Descriptors Based DBLSTM Bottleneck Feature for Speech Driven Talking Avatar

Expressive Speech Driven Talking Avatar Synthesis with DBLSTM Using Limited Amount of Emotional Bimodal Data

Exploring Spatio-Temporal Representations by Integrating Attention-based Bidirectional-LSTM-RNNs and FCNs for Speech Emotion Recognition

Phoneme Embedding and its Application to Speech Driven Talking Avatar Synthesis

A real-time speech driven talking avatar based on deep neural network.

AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation

Auxiliary Features from Laser-Doppler Vibrometer Sensor for Deep Neural Network Based Robust Speech Recognition

Visual Speech Recognition with Lightweight Psychologically Motivated Gabor Features

Acoustic to Articulatory Mapping with Deep Neural Network

Emphasis Detection for Voice Dialogue Applications Using Multi-channel Convolutional Bidirectional Long Short-Term Memory Network

Deep causal speech enhancement and recognition using efficient long-short term memory Recurrent Neural Network

Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar Synthesis

DCIM-AVSR : Efficient Audio-Visual Speech Recognition via Dual Conformer Interaction Module

Look, Listen and Learn - A Multimodal LSTM for Speaker Identification

Sem-Avatar: Semantic Controlled Neural Field for High-Fidelity Audio Driven Avatar.

Auxiliary Multimodal LSTM for Audio-visual Speech Recognition and Lipreading

GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer

LaDTalk: Latent Denoising for Synthesizing Talking Head Videos with High Frequency Details

Learn2Talk: 3D Talking Face Learns from 2D Talking Face

A Novel LSTM-Based Speech Preprocessor for Speaker Diarization in Realistic Mismatch Conditions.

Head and Facial Gestures Synthesis Using PAD Model for an Expressive Talking Avatar