Abstract:This paper proposes Attribute-Decomposed GAN (ADGAN) and its enhanced version (ADGAN++) for controllable image synthesis, which can produce realistic images with desired attributes provided in various source inputs. The core ideas of the proposed ADGAN and ADGAN++ are both to embed component attributes into the latent space as independent codes and thus achieve flexible and continuous control of attributes via mixing and interpolation operations in explicit style representations. The major difference between them is that ADGAN processes all component attributes simultaneously while ADGAN++ utilizes a serial encoding strategy. More specifically, ADGAN consists of two encoding pathways with style block connections and is capable of decomposing the original hard mapping into multiple more accessible subtasks. In the source pathway, component layouts are extracted via a semantic parser and the segmented components are fed into a shared global texture encoder to obtain decomposed latent codes. This strategy allows for the synthesis of more realistic output images and the automatic separation of un-annotated component attributes. Although the original ADGAN works in a delicate and efficient manner, intrinsically it fails to handle the semantic image synthesizing task when the number of attribute categories is huge. To address this problem, ADGAN++ employs the serial encoding of different component attributes to synthesize each part of the target real-world image, and adopts several residual blocks with segmentation guided instance normalization to assemble the synthesized component images and refine the original synthesis result. The two-stage ADGAN++ is designed to alleviate the massive computational costs required when synthesizing real-world images with numerous attributes while maintaining the disentanglement of different attributes to enable flexible control of arbitrary component attributes of the synthesized images. Experimental results demonstrate the proposed methods’ superiority over the state of the art in pose transfer, face style transfer, and semantic image synthesis, as well as their effectiveness in the task of component attribute transfer. Our code and data are publicly available at https://github.com/menyifang/ADGAN .

Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image Generation.

Learn, Imagine and Create: Text-to-Image Generation from Prior Knowledge.

From External to Internal: Structuring Image for Text-to-Image Attributes Manipulation

Diversified text-to-image generation via deep mutual information estimation

DM-GAN: Dynamic Memory Generative Adversarial Networks for Text-to-Image Synthesis

Attribute-Guided Sketch Generation

SAW-GAN: Multi-granularity Text Fusion Generative Adversarial Networks for text-to-image generation

CAGAN: Text-To-Image Generation with Combined Attention GANs

DMF-GAN: Deep Multimodal Fusion Generative Adversarial Networks for Text-to-Image Synthesis

DAE-GAN: Dynamic Aspect-aware GAN for Text-to-Image Synthesis

Object-driven Text-to-Image Synthesis via Adversarial Training

DTGAN: Dual Attention Generative Adversarial Networks for Text-to-Image Generation

Text-to-Image Synthesis via Visual-Memory Creative Adversarial Network.

Controllable Image Synthesis with Attribute-Decomposed GAN

Multi-Sentence Auxiliary Adversarial Networks for Fine-Grained Text-to-Image Synthesis

Adaptive Forgetting, Drafting and Comprehensive Guiding: Text-to-Image Synthesis with Hierarchical Generative Adversarial Networks

KT-GAN: Knowledge-Transfer Generative Adversarial Network for Text-to-Image Synthesis

CF-GAN: cross-domain feature fusion generative adversarial network for text-to-image synthesis

Tf-Gan: Text Feature Fusion Gan for Text-to-Image Generation

Exploring Sparse MoE in GANs for Text-conditioned Image Synthesis.

Multi-Stage Hybrid Text-to-Image Generation Models.