Abstract:Recent advances in text-to-image (T2I) diffusion models have enabled impressive image generation capabilities guided by text prompts. However, extending these techniques to video generation remains challenging, with existing text-to-video (T2V) methods often struggling to produce high-quality and motion-consistent videos. In this work, we introduce Control-A-Video, a controllable T2V diffusion model that can generate videos conditioned on text prompts and reference control maps like edge and depth maps. To tackle video quality and motion consistency issues, we propose novel strategies to incorporate content prior and motion prior into the diffusion-based generation process. Specifically, we employ a first-frame condition scheme to transfer video generation from the image domain. Additionally, we introduce residual-based and optical flow-based noise initialization to infuse motion priors from reference videos, promoting relevance among frame latents for reduced flickering. Furthermore, we present a Spatio-Temporal Reward Feedback Learning (ST-ReFL) algorithm that optimizes the video diffusion model using multiple reward models for video quality and motion consistency, leading to superior outputs. Comprehensive experiments demonstrate that our framework generates higher-quality, more consistent videos compared to existing state-of-the-art methods in controllable text-to-video generation

What problem does this paper attempt to address?

The paper aims to address two main issues in Text-to-Video (T2V) generation: video quality and motion consistency. Specifically, although Text-to-Image (T2I) diffusion models have made significant progress in recent years, generating high-quality images based on text prompts, extending these techniques to video generation remains challenging. Existing T2V methods often struggle to produce high-quality and motion-consistent videos. To tackle these issues, the paper proposes the Control-A-Video model, a controllable T2V diffusion model capable of generating videos based on text prompts and reference control maps (such as edge maps, depth maps, etc.). The main contributions of the paper can be summarized as follows: 1. **Control-A-Video Model**: A controllable T2V diffusion model is proposed, capable of generating videos based on text prompts and reference control maps. 2. **Content Prior**: By utilizing the first frame as a content prior, the model can more easily separate content modeling from temporal modeling during training. A Text-to-Image to Video (T2I-I2V) pipeline is adopted during inference, which helps transfer text-aligned knowledge from images to videos and enables autoregressive generation of longer video sequences. 3. **Motion Prior**: Two innovative noise initialization strategies are introduced—one based on pixel residuals and the other based on optical flow. These motion priors, derived from reference videos, significantly improve inter-frame correlation, resulting in more coherent and less flickering videos. 4. **Spatiotemporal Reward Feedback Learning (ST-ReFL)**: To address quality issues arising during denoising training, the paper proposes a spatiotemporal reward feedback learning algorithm. This algorithm optimizes the trained video diffusion model using various reward models, thereby enhancing video quality and motion consistency. 5. **Experimental Results**: Comprehensive experiments validate that the Control-A-Video model produces higher quality and more consistent videos in controllable text-to-video generation tasks compared to existing state-of-the-art methods. In summary, this paper addresses the issues of video quality and motion consistency in T2V generation by proposing a series of innovative methods and techniques to improve the current state of the art.

Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning

Control-A-Video: Controllable Text-to-Video Generation with Diffusion Models

A Recipe for Scaling Up Text-to-Video Generation with Text-free Videos

ControlVideo: Training-free Controllable Text-to-Video Generation

Controllable Longer Image Animation with Diffusion Models

ART•V: Auto-Regressive Text-to-Video Generation with Diffusion Models

MoVideo: Motion-Aware Video Generation with Diffusion Models

VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNet

Motion-Conditioned Diffusion Model for Controllable Video Synthesis

I4VGen: Image as Free Stepping Stone for Text-to-Video Generation

VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Generative Rendering: Controllable 4D-Guided Video Generation with 2D Diffusion Models

Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling

EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation

Searching Priors Makes Text-to-Video Synthesis Better

TrailBlazer: Trajectory Control for Diffusion-Based Video Generation

HARIVO: Harnessing Text-to-Image Models for Video Generation

InstructVideo: Instructing Video Diffusion Models with Human Feedback

Animate Your Motion: Turning Still Images into Dynamic Videos

Motion-Zero: Zero-Shot Moving Object Control Framework for Diffusion-Based Video Generation