Abstract:The emergence and rapid development of neural networks have been pivotal in advancing text-to-image generative models, with particular emphasis on generative adversarial networks (GANs), variational autoencoders (VAEs), and augmented reality (AR). These models have greatly enriched the field, offering diverse avenues for image generation. Critical support has been provided by databases such as MS COCO, Flickr30K, Visual Genome, and Conceptual Captions, along with essential evaluation metrics, including Inception Score (IS), Frchet Inception Distance (FID), precision, and recall. In this comprehensive review, we delve into the mechanisms and significance of each model and technique, ensuring a holistic examination of their contributions. Both GANs and VAEs stand out as significant models within image generative frameworks, each excelling in distinct aspects. Therefore, it is imperative to discuss both models in this review, as they offer complementary strengths. Additionally, we include noteworthy models such as augmented reality to provide a well-rounded assessment of the current advancements in the field. In terms of datasets, MS COCO offers a diverse and extensive collection of images, serving as a cornerstone for model training. Other datasets like Flickr 30k, Visual Genome, and Conceptual Captions contribute valuable labeled examples, further enriching the learning process for these models. The incorporation of widely recognized metrics and methodologies in the field allows for effective evaluation and comparison of their relative significance. In conclusion, the field's recent achievements owe much to the integration of its various components. VAEs and GANs, with their unique strengths, complement each other, while metrics and datasets play complementary roles in advancing the capabilities of generative models in the context of text-to-image synthesis. This survey underscores the collaborative synergy between models, metrics, and datasets, propelling the field toward new horizons.

Text-to-Image Generation using Generative AI

Mirrorgan: Learning Text-To-Image Generation By Redescription

SIMGAN: Photo-Realistic Semantic Image Manipulation Using Generative Adversarial Networks.

Text-to-Image Synthesis With Generative Models: Methods, Datasets, Performance Metrics, Challenges, and Future Direction

Diversified text-to-image generation via deep mutual information estimation

Text-to-Image Synthesis: A Decade Survey

AI-based text-to-image synthesis: A review

A Survey and Taxonomy of Adversarial Neural Networks for Text-to-Image Synthesis

A Survey of AI Text-to-Image and AI Text-to-Video Generators

RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model

A survey of generative adversarial networks and their application in text-to-image synthesis

Text-to-image Generation Based on Spatial-Channel Attention and Semantic Redescription

Text-to-image Diffusion Models in Generative AI: A Survey

A Comparative Study of Generative Adversarial Networks for Text-to-Image Synthesis

A survey of generative models used in text-to-image

Investigation related to application of Generative Adversarial Networks in text-to-image synthesis

Text-To-Image with Generative Adversarial Networks

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Text2FaceGAN: Face Generation from Fine Grained Textual Descriptions

Generative AI meets 3D: A Survey on Text-to-3D in AIGC Era