MagicStyle: Portrait Stylization Based on Reference Image

Zhaoli Deng,Kaibin Zhou,Fanyi Wang,Zhenpeng Mi

2024-09-12

Abstract:The development of diffusion models has significantly advanced the research on image stylization, particularly in the area of stylizing a content image based on a given style image, which has attracted many scholars. The main challenge in this reference image stylization task lies in how to maintain the details of the content image while incorporating the color and texture features of the style image. This challenge becomes even more pronounced when the content image is a portrait which has complex textural details. To address this challenge, we propose a diffusion model-based reference image stylization method specifically for portraits, called MagicStyle. MagicStyle consists of two phases: Content and Style DDIM Inversion (CSDI) and Feature Fusion Forward (FFF). The CSDI phase involves a reverse denoising process, where DDIM Inversion is performed separately on the content image and the style image, storing the self-attention query, key and value features of both images during the inversion process. The FFF phase executes forward denoising, harmoniously integrating the texture and color information from the pre-stored feature queries, keys and values into the diffusion generation process based on our Well-designed Feature Fusion Attention (FFA). We conducted comprehensive comparative and ablation experiments to validate the effectiveness of our proposed MagicStyle and FFA.

Computer Vision and Pattern Recognition

What problem does this paper attempt to address?

The problem that this paper attempts to solve is: in the style transfer task based on reference images, how to incorporate the texture features of the style image while maintaining the integrity of the details and structure of the content image. This problem is particularly prominent when the content image is a portrait, because portraits usually contain rich details and subtle features, and any improper stylization may lead to image distortion. Specifically, the paper points out: 1. **Challenges**: - How to maintain the details of the content image while introducing the color and texture features of the style image. - When the content image is a portrait, this challenge becomes more complex because the details in the portrait are more abundant and subtle. 2. **Deficiencies of existing methods**: - Existing style transfer methods often have difficulty maintaining the details of the content image and introducing the features of the style image simultaneously when dealing with portraits. To solve these problems, the author proposes a reference - image - stylization method based on the diffusion model, named **MagicStyle**. MagicStyle achieves this goal through two stages: - **Content and Style DDIM Inversion (CSDI)**: Perform DDIM Inversion on the content image and the style image respectively through the reverse denoising process, and store the query, key, and value features in the self - attention mechanism during the process. - **Feature Fusion Forward (FFF)**: Utilize the designed feature - fusion attention mechanism (FFA) to harmoniously integrate the pre - stored feature information into the diffusion generation process, thereby achieving high - quality stylization results. Through this method, MagicStyle can successfully introduce the texture features of the style image while maintaining the details of the content image, providing a new solution for portrait stylization.

MagicStyle: Portrait Stylization Based on Reference Image

Diverse Image Style Transfer Via Invertible Cross-Space Mapping

ArtBank: Artistic Style Transfer with Pre-trained Diffusion Model and Implicit Style Prompt Bank

Image Reference-guided Fashion Design with Structure-aware Transfer by Diffusion Models.

Portrait Diffusion: Training-free Face Stylization with Chain-of-Painting

Artfusion: A Diffusion Model-Based Style Synthesis Framework for Portraits

ZePo: Zero-Shot Portrait Stylization with Faster Sampling

3Dstyle-Diffusion: Pursuing Fine-grained Text-driven 3D Stylization with 2D Diffusion Models

ControlStyle: Text-Driven Stylized Image Generation Using Diffusion Priors

Selective Image Abstraction

DiffStyler: Diffusion-based Localized Image Style Transfer

InstaStyle: Inversion Noise of a Stylized Image is Secretly a Style Adviser

Style3D: Attention-guided Multi-view Style Transfer for 3D Object Generation

Style Image Harmonization Via Global-Local Style Mutual Guided

Towards Multi-View Consistent Style Transfer with One-Step Diffusion via Vision Conditioning

Artistic Style Transfer Based on Attention with Knowledge Distillation

Artistic stylization of face photos based on a single exemplar

ArtFusion: Controllable Arbitrary Style Transfer using Dual Conditional Latent Diffusion Models

FreeStyle: Free Lunch for Text-guided Style Transfer using Diffusion Models

Magic Insert: Style-Aware Drag-and-Drop

InstantStyle: Free Lunch towards Style-Preserving in Text-to-Image Generation