Abstract:Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head video clips which simultaneously have accurate lip synchronization and motion smoothness. Previous approaches, including 3DMM-based (3D Morphable Model) methods and NeRF-based (Neural Radiance Field) methods, are sub-optimal in that they either require minutes of source videos and days of training time or lack the disentangled control of verbal (e.g., lip motion) and non-verbal (e.g., head pose and expression) representations for video clip insertion. In this work, we fully utilize the video context to design a novel framework for talking-head video editing, which achieves efficiency, disentangled motion control, and sequential smoothness. Specifically, we decompose this framework to motion prediction and motion-conditioned rendering: (1) We first design an animation prediction module that efficiently obtains smooth and lip-sync motion sequences conditioned on the driven speech. This module adopts a non-autoregressive network to obtain context prior and improve the prediction efficiency, and it learns a speech-animation mapping prior with better generalization to novel speech from a multi-identity video dataset. (2) We then introduce a neural rendering module to synthesize the photo-realistic and full-head video frames given the predicted motion sequence. This module adopts a pre-trained head topology and uses only few frames for efficient fine-tuning to obtain a person-specific rendering model. Extensive experiments demonstrate that our method efficiently achieves smoother editing results with higher image quality and lip accuracy using less data than previous methods.

Editing like Humans: A Contextual, Multimodal Framework for Automated Video Editing

ExpressEdit: Video Editing with Natural Language and Sketching

Edit As You Wish: Video Caption Editing with Multi-grained User Control

FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing

The Anatomy of Video Editing: A Dataset and Benchmark Suite for AI-Assisted Video Editing

VCoME: Verbal Video Composition with Multimodal Editing Effects

M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers

VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing

Context-Aware Talking-Head Video Editing

A Reinforcement Learning-Based Automatic Video Editing Method Using Pre-trained Vision-Language Model

Automatic Non-Linear Video Editing Transfer

Motion-Conditioned Image Animation for Video Editing

Video Editing for Video Retrieval

Temporally Consistent Object Editing in Videos using Extended Attention

Text-based editing of talking-head video

Story-driven Video Editing

EditBoard: Towards A Comprehensive Evaluation Benchmark for Text-based Video Editing Models

Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

Pathways on the Image Manifold: Image Editing via Video Generation

UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing

CCEdit: Creative and Controllable Video Editing via Diffusion Models