Abstract:We consider the problem of independently, in a disentangled fashion, controlling the outputs of text-to-image diffusion models with color and style attributes of a user-supplied reference image. We present the first training-free, test-time-only method to disentangle and condition text-to-image models on color and style attributes from reference image. To realize this, we propose two key innovations. Our first contribution is to transform the latent codes at inference time using feature transformations that make the covariance matrix of current generation follow that of the reference image, helping meaningfully transfer color. Next, we observe that there exists a natural disentanglement between color and style in the LAB image space, which we exploit to transform the self-attention feature maps of the image being generated with respect to those of the reference computed from its L channel. Both these operations happen purely at test time and can be done independently or merged. This results in a flexible method where color and style information can come from the same reference image or two different sources, and a new generation can seamlessly fuse them in either scenario.

What problem does this paper attempt to address?

### What problems does this paper attempt to solve? This paper aims to solve the problem of how to independently control the color and style attributes of the generated image without additional training in the text - to - image generation model. Specifically, the author hopes to achieve the following goals: 1. **Independent control of color and style**: Extract the color and style attributes from the reference image provided by the user and be able to independently apply these attributes when generating a new image. This means that the color information and style information can come from different sources respectively, and these information can be flexibly combined during the generation process. 2. **Completely no - training - required**: The proposed method should be able to be directly used at the test time without additional training or fine - tuning of the model. This makes the method more practical because there is no need to retrain the model whenever the reference image changes. 3. **Multi - attribute fusion**: Allow multiple attributes (such as color and style) to be fused simultaneously in one generation process, and these attributes can come from the same reference image or two different reference images. To achieve the above goals, the author proposes two innovative methods: - **Time - step - limited latent code recoloring transformation**: Align the color of the generated image with that of the reference image by adjusting the covariance matrix of the latent code. - **Self - attention feature operation in LAB space**: Utilize the natural color and style separation characteristics of the L channel (luminance channel) in the LAB color space, and perform feature operations only at specific time steps during the generation process to achieve style transfer. These two methods work together to achieve decoupled control of color and style in the text - to - image generation process.

Training-free Color-Style Disentanglement for Constrained Text-to-Image Synthesis

UATST: Towards Unpaired Arbitrary Text-Guided Style Transfer with Cross-Space Modulation

Style Permutation for Diversified Arbitrary Style Transfer

Diverse Image Style Transfer Via Invertible Cross-Space Mapping

TeSTNeRF: Text-Driven 3D Style Transfer Via Cross-Modal Learning.

StyleDiffusion: Controllable Disentangled Style Transfer via Diffusion Models

ColorEdit: Training-free Image-Guided Color editing with diffusion model

Test-time Conditional Text-to-Image Synthesis Using Diffusion Models

ColorwAI: Generative Colorways of Textiles through GAN and Diffusion Disentanglement

Style Transformer for Image Inversion and Editing

PreciseControl: Enhancing Text-To-Image Diffusion Models with Fine-Grained Attribute Control

FreeStyle: Free Lunch for Text-guided Style Transfer using Diffusion Models

Towards Spatially Disentangled Manipulation of Face Images With Pre-Trained StyleGANs

StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements

DiffColor: Toward High Fidelity Text-Guided Image Colorization with Diffusion Models

Training-free Composite Scene Generation for Layout-to-Image Synthesis

Training-Free Sketch-Guided Diffusion with Latent Optimization

ControlStyle: Text-Driven Stylized Image Generation Using Diffusion Priors

Arbitrary Style Guidance for Enhanced Diffusion-Based Text-to-Image Generation

Disentangling for Text-to-Image Generation