Abstract:The emergence and rapid development of neural networks have been pivotal in advancing text-to-image generative models, with particular emphasis on generative adversarial networks (GANs), variational autoencoders (VAEs), and augmented reality (AR). These models have greatly enriched the field, offering diverse avenues for image generation. Critical support has been provided by databases such as MS COCO, Flickr30K, Visual Genome, and Conceptual Captions, along with essential evaluation metrics, including Inception Score (IS), Frchet Inception Distance (FID), precision, and recall. In this comprehensive review, we delve into the mechanisms and significance of each model and technique, ensuring a holistic examination of their contributions. Both GANs and VAEs stand out as significant models within image generative frameworks, each excelling in distinct aspects. Therefore, it is imperative to discuss both models in this review, as they offer complementary strengths. Additionally, we include noteworthy models such as augmented reality to provide a well-rounded assessment of the current advancements in the field. In terms of datasets, MS COCO offers a diverse and extensive collection of images, serving as a cornerstone for model training. Other datasets like Flickr 30k, Visual Genome, and Conceptual Captions contribute valuable labeled examples, further enriching the learning process for these models. The incorporation of widely recognized metrics and methodologies in the field allows for effective evaluation and comparison of their relative significance. In conclusion, the field's recent achievements owe much to the integration of its various components. VAEs and GANs, with their unique strengths, complement each other, while metrics and datasets play complementary roles in advancing the capabilities of generative models in the context of text-to-image synthesis. This survey underscores the collaborative synergy between models, metrics, and datasets, propelling the field toward new horizons.

A survey of generative models used in text-to-image

Learn, Imagine and Create: Text-to-Image Generation from Prior Knowledge.

Text-to-Image Synthesis With Generative Models: Methods, Datasets, Performance Metrics, Challenges, and Future Direction

A survey of generative adversarial networks and their application in text-to-image synthesis

Investigation related to application of Generative Adversarial Networks in text-to-image synthesis

A Survey and Taxonomy of Adversarial Neural Networks for Text-to-Image Synthesis

Generative Adversarial Networks in Computer Vision: A Survey and Taxonomy

RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model

A survey on GANs for computer vision: Recent research, analysis and taxonomy

Image Synthesis with Adversarial Networks: a Comprehensive Survey and Case Studies

Comparison and Analysis of Image-to-Image Generative Adversarial Networks: A Survey

Recent Advances of Generative Adversarial Networks in Computer Vision

Recent Advance On Generative Adversarial Networks

A Survey of AI Text-to-Image and AI Text-to-Video Generators

An introduction to image synthesis with generative adversarial nets

P‐2.9: A review of image generation methods based on deep learning

Comparative Analysis of Generative Models: Enhancing Image Synthesis with VAEs, GANs, and Stable Diffusion

Review of Research on Improvement and Application of Generative Adversarial Networks

A State-of-the-Art Review on Image Synthesis With Generative Adversarial Networks

Survey on Visual Signal Coding and Processing with Generative Models: Technologies, Standards and Optimization

Survey on Generative Adversarial Behavior in Artificial Neural Tasks