Abstract:This paper proposes Attribute-Decomposed GAN (ADGAN) and its enhanced version (ADGAN++) for controllable image synthesis, which can produce realistic images with desired attributes provided in various source inputs. The core ideas of the proposed ADGAN and ADGAN++ are both to embed component attributes into the latent space as independent codes and thus achieve flexible and continuous control of attributes via mixing and interpolation operations in explicit style representations. The major difference between them is that ADGAN processes all component attributes simultaneously while ADGAN++ utilizes a serial encoding strategy. More specifically, ADGAN consists of two encoding pathways with style block connections and is capable of decomposing the original hard mapping into multiple more accessible subtasks. In the source pathway, component layouts are extracted via a semantic parser and the segmented components are fed into a shared global texture encoder to obtain decomposed latent codes. This strategy allows for the synthesis of more realistic output images and the automatic separation of un-annotated component attributes. Although the original ADGAN works in a delicate and efficient manner, intrinsically it fails to handle the semantic image synthesizing task when the number of attribute categories is huge. To address this problem, ADGAN++ employs the serial encoding of different component attributes to synthesize each part of the target real-world image, and adopts several residual blocks with segmentation guided instance normalization to assemble the synthesized component images and refine the original synthesis result. The two-stage ADGAN++ is designed to alleviate the massive computational costs required when synthesizing real-world images with numerous attributes while maintaining the disentanglement of different attributes to enable flexible control of arbitrary component attributes of the synthesized images. Experimental results demonstrate the proposed methods’ superiority over the state of the art in pose transfer, face style transfer, and semantic image synthesis, as well as their effectiveness in the task of component attribute transfer. Our code and data are publicly available at https://github.com/menyifang/ADGAN .

Controllable Image Synthesis with Attribute-Decomposed GAN

Fine-grained Semantic Constraint in Image Synthesis

Style Fader Generative Adversarial Networks for Style Degree Controllable Artistic Style Transfer

Attribute-Guided Sketch Generation

Creative and Diverse Artwork Generation Using Adversarial Networks

Customizable GAN: Customizable Image Synthesis Based on Adversarial Learning.

Fully-Featured Attribute Transfer

Attribute-specific Control Units in StyleGAN for Fine-grained Image Manipulation

Two Birds with One Stone: Transforming and Generating Facial Images with Iterative GAN

Two Birds with One Stone: Iteratively Learn Facial Attributes with GANs.

Adma-GAN: Attribute-Driven Memory Augmented GANs for Text-to-Image Generation.

Pose- and Attribute-consistent Person Image Synthesis

Dual Attention GANs for Semantic Image Synthesis

Controllable and Identity-Aware Facial Attribute Transformation

Core-attributes enhanced generative adversarial networks for robust image enhancement

Controllable facial attribute editing via Gaussian mixture model disentanglement

InDecGAN: Learning to Generate Complex Images from Captions Via Independent Object-Level Decomposition and Enhancement

Prominent Attribute Modification using Attribute Dependent Generative Adversarial Network

Generative Semantic Manipulation with Contrasting GAN

TAGE: Trustworthy Attribute Group Editing for Stable Few-shot Image Generation

GL-GAN: Adaptive Global and Local Bilevel Optimization model of Image Generation