Adam Polyak,Amit Zohar,Andrew Brown,Andros Tjandra,Animesh Sinha,Ann Lee,Apoorv Vyas,Bowen Shi,Chih-Yao Ma,Ching-Yao Chuang,David Yan,Dhruv Choudhary,Dingkang Wang,Geet Sethi,Guan Pang,Haoyu Ma,Ishan Misra,Ji Hou,Jialiang Wang,Kiran Jagadeesh,Kunpeng Li,Luxin Zhang,Mannat Singh,Mary Williamson,Matt Le,Matthew Yu,Mitesh Kumar Singh,Peizhao Zhang,Peter Vajda,Quentin Duval,Rohit Girdhar,Roshan Sumbaly,Sai Saketh Rambhatla,Sam Tsai,Samaneh Azadi,Samyak Datta,Sanyuan Chen,Sean Bell,Sharadh Ramaswamy,Shelly Sheynin,Siddharth Bhattacharya,Simran Motwani,Tao Xu,Tianhe Li,Tingbo Hou,Wei-Ning Hsu,Xi Yin,Xiaoliang Dai,Yaniv Taigman,Yaqiao Luo,Yen-Cheng Liu,Yi-Chiao Wu,Yue Zhao,Yuval Kirstain,Zecheng He,Zijian He,Albert Pumarola,Ali Thabet,Artsiom Sanakoyeu,Arun Mallya,Baishan Guo,Boris Araya,Breena Kerr,Carleigh Wood,Ce Liu,Cen Peng,Dimitry Vengertsev,Edgar Schonfeld,Elliot Blanchard,Felix Juefei-Xu,Fraylie Nord,Jeff Liang,John Hoffman,Jonas Kohler,Kaolin Fire,Karthik Sivakumar,Lawrence Chen,Licheng Yu,Luya Gao,Markos Georgopoulos,Rashel Moritz,Sara K. Sampson,Shikai Li,Simone Parmeggiani,Steve Fine,Tara Fowler,Vladan Petrovic,Yuming Du

Abstract:We present Movie Gen, a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. We also show additional capabilities such as precise instruction-based video editing and generation of personalized videos based on a user's image. Our models set a new state-of-the-art on multiple tasks: text-to-video synthesis, video personalization, video editing, video-to-audio generation, and text-to-audio generation. Our largest video generation model is a 30B parameter transformer trained with a maximum context length of 73K video tokens, corresponding to a generated video of 16 seconds at 16 frames-per-second. We show multiple technical innovations and simplifications on the architecture, latent spaces, training objectives and recipes, data curation, evaluation protocols, parallelization techniques, and inference optimizations that allow us to reap the benefits of scaling pre-training data, model size, and training compute for training large scale media generation models. We hope this paper helps the research community to accelerate progress and innovation in media generation models. All videos from this paper are available at <a class="link-external link-https" href="https://go.fb.me/MovieGenResearchVideos" rel="external noopener nofollow">this https URL</a>.

What problem does this paper attempt to address?

The problem that this paper attempts to solve is to generate high - quality, multi - purpose multimedia content, specifically including: 1. **High - quality video generation**: Generate high - resolution (1080p HD) and synchronous audio - video with different aspect ratios. 2. **Personalized video generation**: Generate personalized video content based on the user's image. 3. **Precise video editing**: Perform precise video editing based on user instructions. 4. **Text - to - video synthesis**: Generate videos from text prompts. 5. **Video - to - audio generation**: Generate synchronous audio according to the video. 6. **Text - to - audio generation**: Generate audio according to text prompts. ### Detailed problem description - **High - quality video generation**: Existing video generation models face challenges when generating high - quality, long - duration, multi - resolution videos. Movie Gen has solved these problems through large - scale pre - training and fine - tuning, and is able to generate high - definition videos up to 16 seconds long. - **Personalized video generation**: Current video generation models have difficulty generating personalized videos based on the images of specific individuals. Movie Gen has solved this problem by introducing the Personalized Movie Gen Video model, and is able to maintain the identity characteristics of individuals in the video. - **Precise video editing**: Existing video editing tools usually require a large amount of supervised data and are difficult to perform complex editing tasks. Movie Gen has achieved precise video editing based on text instructions by proposing a training method without supervised data. - **Text - to - video synthesis**: The text - to - video task requires the model to understand the text and generate high - quality videos that match it. Movie Gen has improved the quality of text - to - video generation by using techniques such as Flow Matching. - **Video - to - audio generation**: Audio generation in videos requires the model to understand the visual content and generate appropriate audio in synchronization with it. The Movie Gen Audio model is able to generate realistic sound effects and music according to the video content. - **Text - to - audio generation**: The text - to - audio task requires the model to generate natural and smooth audio according to text prompts. Movie Gen Audio has improved the quality of text - to - audio generation through large - scale pre - training and fine - tuning. ### Technological innovation To achieve these goals, the paper proposes several technological innovations: - **Transformer architecture**: Use a Transformer model with 30B parameters for large - scale pre - training to generate high - quality videos and audio. - **Flow Matching training objective**: Adopt the Flow Matching framework to improve the generation quality and ensure that the generated content is more realistic. - **Spatiotemporal - compressed latent space**: Compress videos into a low - dimensional latent space through Temporal Autoencoder (TAE), thereby reducing the amount of computation and improving the generation efficiency. - **Rich text embedding**: Combine multiple pre - trained text encoders to provide multi - level text understanding capabilities. - **Spatial upsampling**: Convert low - resolution videos into high - resolution videos through Spatial Upsampler to reduce the computational cost. Through these technologies and methods, Movie Gen has reached a new state - of - the - art level in multiple media generation tasks and provides valuable benchmarks and references for future research.

Movie Gen: A Cast of Media Foundation Models

HunyuanVideo: A Systematic Framework For Large Video Generative Models

Movie Gen: SWOT Analysis of Meta's Generative AI Foundation Model for Transforming Media Generation, Advertising, and Entertainment Industries

MovieFactory: Automatic Movie Creation from Text using Large Generative Models for Language and Images

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Imagen Video: High Definition Video Generation with Diffusion Models

xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations

EvalCrafter: Benchmarking and Evaluating Large Video Generation Models

MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence

Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising

MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation

UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation

MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

MovieLLM: Enhancing Long Video Understanding with AI-Generated Movies

VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

VideoPoet: A Large Language Model for Zero-Shot Video Generation

Latent Video Diffusion Models for High-Fidelity Long Video Generation

Towards Smooth Video Composition

MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation

The Dawn of Video Generation: Preliminary Explorations with SORA-like Models