Abstract:The emergence and rapid development of neural networks have been pivotal in advancing text-to-image generative models, with particular emphasis on generative adversarial networks (GANs), variational autoencoders (VAEs), and augmented reality (AR). These models have greatly enriched the field, offering diverse avenues for image generation. Critical support has been provided by databases such as MS COCO, Flickr30K, Visual Genome, and Conceptual Captions, along with essential evaluation metrics, including Inception Score (IS), Frchet Inception Distance (FID), precision, and recall. In this comprehensive review, we delve into the mechanisms and significance of each model and technique, ensuring a holistic examination of their contributions. Both GANs and VAEs stand out as significant models within image generative frameworks, each excelling in distinct aspects. Therefore, it is imperative to discuss both models in this review, as they offer complementary strengths. Additionally, we include noteworthy models such as augmented reality to provide a well-rounded assessment of the current advancements in the field. In terms of datasets, MS COCO offers a diverse and extensive collection of images, serving as a cornerstone for model training. Other datasets like Flickr 30k, Visual Genome, and Conceptual Captions contribute valuable labeled examples, further enriching the learning process for these models. The incorporation of widely recognized metrics and methodologies in the field allows for effective evaluation and comparison of their relative significance. In conclusion, the field's recent achievements owe much to the integration of its various components. VAEs and GANs, with their unique strengths, complement each other, while metrics and datasets play complementary roles in advancing the capabilities of generative models in the context of text-to-image synthesis. This survey underscores the collaborative synergy between models, metrics, and datasets, propelling the field toward new horizons.

A survey of text generation models

Neural Text Generation: Past, Present and Beyond

The survey: Text generation models in deep learning

A survey of generative models used in text-to-image

A Systematic survey on automated text generation tools and techniques: application, evaluation, and challenges

Pre-trained Language Models for Text Generation: A Survey

A survey on text generation using generative adversarial networks

Pretrained Language Models for Text Generation: A Survey

Diffusion models in text generation: a survey

Diffusion Models for Non-autoregressive Text Generation: A Survey

An Overview on Controllable Text Generation via Variational Auto-Encoders

A Survey of Data-Driven 2D Diffusion Models for Generating Images from Text

Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges

Which Discriminator for Cooperative Text Generation?

Recent Advances in Neural Text Generation: A Task-Agnostic Survey

Text Generation Based on Generative Adversarial Nets with Latent Variable

A Survey On Text-to-3D Contents Generation In The Wild

A Reparameterized Discrete Diffusion Model for Text Generation

A Combined Encoder and Transformer Approach for Coherent and High-Quality Text Generation

Natural Language Generation Using Sequential Models: A Survey

Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework