Abstract:In this paper, we study an object synthesis task that combines an object text with an object image to create a new object image. However, most diffusion models struggle with this task, \textit{i.e.}, often generating an object that predominantly reflects either the text or the image due to an imbalance between their inputs. To address this issue, we propose a simple yet effective method called Adaptive Text-Image Harmony (ATIH) to generate novel and surprising objects. First, we introduce a scale factor and an injection step to balance text and image features in cross-attention and to preserve image information in self-attention during the text-image inversion diffusion process, respectively. Second, to better integrate object text and image, we design a balanced loss function with a noise parameter, ensuring both optimal editability and fidelity of the object image. Third, to adaptively adjust these parameters, we present a novel similarity score function that not only maximizes the similarities between the generated object image and the input text/image but also balances these similarities to harmonize text and image integration. Extensive experiments demonstrate the effectiveness of our approach, showcasing remarkable object creations such as colobus-glass jar. Project page: <a class="link-external link-https" href="https://xzr52.github.io/ATIH/" rel="external noopener nofollow">this https URL</a>.

What problem does this paper attempt to address?

The problem that this paper attempts to solve is: when generating new object images, how to balance the imbalance between text and image inputs. Specifically, existing diffusion models (such as SDXL - Turbo) often tend to reflect one of the text or image when combining text and image to generate new objects, resulting in poor fusion effects. For example, when given an animal image and a text description of a different category, the generated new object image may only reflect the text content or the original image, and fail to effectively combine the features of both. To solve this problem, the author proposes a method named **Adaptive Text - Image Harmony (ATIH)**. This method balances the weights of text and image features in cross - attention and self - attention by introducing the scale factor \(\alpha\) and the injection step \(i\), and designs a balanced loss function to optimize the noise parameters, thereby ensuring that the generated new object image can retain the features of both text and image at the same time and achieve a more harmonious fusion. The following are the main contributions of the ATIH method: 1. **Introducing scale factor and injection step**: In the reverse diffusion process, the scale factor \(\alpha\) and the injection step \(i\) are introduced to balance the weights of text and image features in cross - attention and retain image information in self - attention. 2. **Designing a balanced loss function**: By designing a balanced loss function containing noise parameters, it is ensured that the generated image achieves the best balance between reconstruction and Gaussian white noise approximation, thereby improving editability and fidelity. 3. **Proposing a novel similarity scoring function**: This function not only maximizes the similarity between the generated image and the input text / image, but also balances these similarities by adjusting the scale factor and the injection step to achieve harmonious fusion of text and image. Through these improvements, the ATIH method can better combine the features of text and image when generating new object images, and produce more creative and high - quality composite images. Experimental results show that this method is superior to existing image editing and creative mixing methods on multiple datasets.

Novel Object Synthesis via Adaptive Text-Image Harmony

Novel 3D-Aware Composition Images Synthesis for Object Display with Diffusion Model.

Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image Synthesis

CreativeSynth: Creative Blending and Synthesis of Visual Arts based on Multimodal Diffusion

Object-driven Text-to-Image Synthesis via Adversarial Training

Object-Driven One-Shot Fine-tuning of Text-to-Image Diffusion with Prototypical Embedding

Improving Text-guided Object Inpainting with Semantic Pre-inpainting

UDiffText: A Unified Framework for High-quality Text Synthesis in Arbitrary Images via Character-aware Diffusion Models

LTOS: Layout-controllable Text-Object Synthesis via Adaptive Cross-attention Fusions

Customizable GAN: Customizable Image Synthesis Based on Adversarial Learning.

eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Enhancing Object Coherence in Layout-to-Image Synthesis

TP2O: Creative Text Pair-to-Object Generation using Balance Swap-Sampling

DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation

Training-free Composite Scene Generation for Layout-to-Image Synthesis

Scene Text Synthesis for Efficient and Effective Deep Network Training

Brush Your Text: Synthesize Any Scene Text on Images via Diffusion Model

Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models

Harmonizing Visual Text Comprehension and Generation

Text-image Alignment for Diffusion-based Perception

Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation