Abstract:The rapid advancement of text-to-image (T2I) diffusion models has enabled them to generate unprecedented results from given texts. However, as text inputs become longer, existing encoding methods like CLIP face limitations, and aligning the generated images with long texts becomes challenging. To tackle these issues, we propose LongAlign, which includes a segment-level encoding method for processing long texts and a decomposed preference optimization method for effective alignment training. For segment-level encoding, long texts are divided into multiple segments and processed separately. This method overcomes the maximum input length limits of pretrained encoding models. For preference optimization, we provide decomposed CLIP-based preference models to fine-tune diffusion models. Specifically, to utilize CLIP-based preference models for T2I alignment, we delve into their scoring mechanisms and find that the preference scores can be decomposed into two components: a text-relevant part that measures T2I alignment and a text-irrelevant part that assesses other visual aspects of human preference. Additionally, we find that the text-irrelevant part contributes to a common overfitting problem during fine-tuning. To address this, we propose a reweighting strategy that assigns different weights to these two components, thereby reducing overfitting and enhancing alignment. After fine-tuning $512 \times 512$ Stable Diffusion (SD) v1.5 for about 20 hours using our method, the fine-tuned SD outperforms stronger foundation models in T2I alignment, such as PixArt-$\alpha$ and Kandinsky v2.2. The code is available at <a class="link-external link-https" href="https://github.com/luping-liu/LongAlign" rel="external noopener nofollow">this https URL</a>.

Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences through $f$-divergence Minimization

$f$-PO: Generalizing Preference Optimization with $f$-divergence Minimization

Margin-aware Preference Optimization for Aligning Diffusion Models without Reference

Diff-Instruct++: Training One-step Text-to-image Generator Model to Align with Human Preferences

Improving Long-Text Alignment for Text-to-Image Diffusion Models

A Dense Reward View on Aligning Text-to-Image Diffusion with Preference

Diffusion-RPO: Aligning Diffusion Models through Relative Preference Optimization

Aligning Text-to-Image Diffusion Models with Reward Backpropagation

Aligning Diffusion Models by Optimizing Human Utility

Stable Preference: Redefining Training Paradigm of Human Preference Model for Text-to-Image Synthesis

On the Relation Between Quality-Diversity Evaluation and Distribution-Fitting Goal in Text Generation

Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment

Direct Preference Optimization Using Sparse Feature-Level Constraints

Diversity-Promoting GAN: A Cross-Entropy Based Generative Adversarial Network for Diversified Text Generation

Exploring the Pareto-Optimality between Quality and Diversity in Text Generation

Diffusion Model Alignment Using Direct Preference Optimization

Bridging the Gap: Aligning Text-to-Image Diffusion Models with Specific Feedback

Towards Better Text-to-Image Generation Alignment via Attention Modulation

Discriminative Probing and Tuning for Text-to-Image Generation

Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization