Abstract:Personalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained identity details of the generated results. Furthermore, there is little prior work tackling this problem. To help improve these identity details in the personalized image generation, we introduce a new task: reference-guided artifacts refinement. We present Refine-by-Align, a first-of-its-kind model that employs a diffusion-based framework to address this challenge. Our model consists of two stages: Alignment Stage and Refinement Stage, which share weights of a unified neural network model. Given a generated image, a masked artifact region, and a reference image, the alignment stage identifies and extracts the corresponding regional features in the reference, which are then used by the refinement stage to fix the artifacts. Our model-agnostic pipeline requires no test-time tuning or optimization. It automatically enhances image fidelity and reference identity in the generated image, generalizing well to existing models on various tasks including but not limited to customization, generative compositing, view synthesis, and virtual try-on. Extensive experiments and comparisons demonstrate that our pipeline greatly pushes the boundary of fine details in the image synthesis models.

What problem does this paper attempt to address?

The problem that this paper attempts to solve is that in personalized image generation, the generated images often have local flaws (such as incorrect logos or fonts), which reduce the realism of the images and the fineness of details. Although the existing generation models have made remarkable progress in personalized image generation, local flaws often appear in the images generated by these models, such as incorrect logos, font errors, etc., and these problems affect the fidelity of the generated results and the fine - grained identity details. In addition, there are relatively few research works on this problem. Therefore, this paper proposes a new task: reference - guided artifacts refinement, aiming to use the relevant information in the reference image to improve these flaws in the generated image, thereby improving the quality and fidelity of the generated image. Specifically, the paper introduces a new model named "Refine - by - Align". This model adopts a diffusion framework to address the above challenges. The model is divided into two stages: the Alignment Stage and the Refinement Stage. In the Alignment Stage, the model guides the Refinement Stage to repair the flaws in the generated image by identifying and extracting the corresponding regional features in the reference image. This method does not require test - time tuning or optimization, can automatically enhance the realism of the generated image and the ability to maintain the reference identity, and can be well generalized to various tasks of existing models, such as customization, generation combination, view synthesis and virtual try - on, etc. In summary, the main contribution of this paper is to provide a generation flaw repair framework that can be controlled by specifying a reference image. This framework can handle flaws of any shape and size while maintaining the identity and background information of the original generated image.

Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment

FaceChain: A Playground for Identity-Preserving Portrait Generation

G-Refine: A General Quality Refiner for Text-to-Image Generation

RealisHuman: A Two-Stage Approach for Refining Malformed Human Parts in Generated Images

HumanRefiner: Benchmarking Abnormal Human Generation and Refining with Coarse-to-fine Pose-Reversible Guidance

RefFaceNet: Reference-based Face Image Generation from Line Art Drawings

Q-Refine: A Perceptual Quality Refiner for AI-Generated Image

Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image Generation

Two Birds with One Stone: Transforming and Generating Facial Images with Iterative GAN

PersonaCraft: Personalized Full-Body Image Synthesis for Multiple Identities from Single References Using 3D-Model-Conditioned Diffusion

Refining CycleGAN with attention mechanisms and age-Aware training for realistic Deepfakes

Fine-grained Identity Preserving Landmark Synthesis for Face Reenactment

Imagine yourself: Tuning-Free Personalized Image Generation

RefDrop: Controllable Consistency in Image or Video Generation via Reference Feature Guidance

ID-Aligner: Enhancing Identity-Preserving Text-to-Image Generation with Reward Feedback Learning

HumanRef: Single Image to 3D Human Generation via Reference-Guided Diffusion

Rendering-Refined Stable Diffusion for Privacy Compliant Synthetic Data

DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation

HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting

SpecRef: A Fast Training-free Baseline of Specific Reference-Condition Real Image Editing

SynArtifact: Classifying and Alleviating Artifacts in Synthetic Images via Vision-Language Model