Abstract:Text-to-image generation intends to automatically produce a photo-realistic image, conditioned on a textual description. It can be potentially employed in the field of art creation, data augmentation, photo-editing, etc. Although many efforts have been dedicated to this task, it remains particularly challenging to generate believable, natural scenes. To facilitate the real-world applications of text-to-image synthesis, we focus on studying the following three issues: 1) How to ensure that generated samples are believable, realistic or natural? 2) How to exploit the latent space of the generator to edit a synthesized image? 3) How to improve the explainability of a text-to-image generation framework? In this work, we constructed two novel data sets (i.e., the Good & Bad bird and face data sets) consisting of successful as well as unsuccessful generated samples, according to strict criteria. To effectively and efficiently acquire high-quality images by increasing the probability of generating Good latent codes, we use a dedicated Good/Bad classifier for generated images. It is based on a pre-trained front end and fine-tuned on the basis of the proposed Good & Bad data set. After that, we present a novel algorithm which identifies semantically-understandable directions in the latent space of a conditional text-to-image GAN architecture by performing independent component analysis on the pre-trained weight values of the generator. Furthermore, we develop a background-flattening loss (BFL), to improve the background appearance in the edited image. Subsequently, we introduce linear interpolation analysis between pairs of keywords. This is extended into a similar triangular `linguistic' interpolation in order to take a deep look into what a text-to-image synthesis model has learned within the linguistic embeddings. Our data set is available at <a class="link-external link-https" href="https://zenodo.org/record/6283798#.YhkN_ujMI2w" rel="external noopener nofollow">this https URL</a>.

IRC-GAN: Introspective Recurrent Convolutional GAN for Text-to-video Generation.

Learn, Imagine and Create: Text-to-Image Generation from Prior Knowledge.

A Recipe for Scaling Up Text-to-Video Generation with Text-free Videos

Diversified text-to-image generation via deep mutual information estimation

To Create What You Tell: Generating Videos from Captions

Scripted Video Generation With a Bottom-Up Generative Adversarial Network

Recurrent Deconvolutional Generative Adversarial Networks with Application to Text Guided Video Generation

ISF-GAN: Imagine, Select, and Fuse with GPT-Based Text Enrichment for Text-to-Image Synthesis

Exploring AIGC Video Quality: A Focus on Visual Harmony, Video-Text Consistency and Domain Distribution Gap

OptGAN: Optimizing and Interpreting the Latent Space of the Conditional Text-to-Image GANs

RIGID: Recurrent GAN Inversion and Editing of Real Face Videos

Facial Expression Video Generation Based-On Spatio-temporal Convolutional GAN: FEV-GAN

Attention-Based Image-to-Video Translation for Synthesizing Facial Expression Using GAN

Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation

Automated Visual Generation using GAN with Textual Information Feeds

xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations

KT-GAN: Knowledge-Transfer Generative Adversarial Network for Text-to-Image Synthesis

CgT-GAN: CLIP-guided Text GAN for Image Captioning

CT-GAN: A conditional Generative Adversarial Network of transformer architecture for text-to-image

DF-GAN: Deep Fusion Generative Adversarial Networks for Text-to-Image Synthesis.