Abstract:The emergence and rapid development of neural networks have been pivotal in advancing text-to-image generative models, with particular emphasis on generative adversarial networks (GANs), variational autoencoders (VAEs), and augmented reality (AR). These models have greatly enriched the field, offering diverse avenues for image generation. Critical support has been provided by databases such as MS COCO, Flickr30K, Visual Genome, and Conceptual Captions, along with essential evaluation metrics, including Inception Score (IS), Frchet Inception Distance (FID), precision, and recall. In this comprehensive review, we delve into the mechanisms and significance of each model and technique, ensuring a holistic examination of their contributions. Both GANs and VAEs stand out as significant models within image generative frameworks, each excelling in distinct aspects. Therefore, it is imperative to discuss both models in this review, as they offer complementary strengths. Additionally, we include noteworthy models such as augmented reality to provide a well-rounded assessment of the current advancements in the field. In terms of datasets, MS COCO offers a diverse and extensive collection of images, serving as a cornerstone for model training. Other datasets like Flickr 30k, Visual Genome, and Conceptual Captions contribute valuable labeled examples, further enriching the learning process for these models. The incorporation of widely recognized metrics and methodologies in the field allows for effective evaluation and comparison of their relative significance. In conclusion, the field's recent achievements owe much to the integration of its various components. VAEs and GANs, with their unique strengths, complement each other, while metrics and datasets play complementary roles in advancing the capabilities of generative models in the context of text-to-image synthesis. This survey underscores the collaborative synergy between models, metrics, and datasets, propelling the field toward new horizons.

A review of multimodal learning for text to images

Learn, Imagine and Create: Text-to-Image Generation from Prior Knowledge.

Diversified text-to-image generation via deep mutual information estimation

A Review of Multi-Modal Learning from the Text-Guided Visual Processing Viewpoint

A survey of generative adversarial networks and their application in text-to-image synthesis

Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models

Text-to-Image Cross-Modal Generation: A Systematic Review

A survey of generative models used in text-to-image

Adaptive multi-text union for stable text-to-image synthesis learning

Unified Text-to-Image Generation and Retrieval

A Survey on Image-text Multimodal Models

Multi-Modal Generative AI: Multi-modal LLM, Diffusion and Beyond

Text-to-image Generation Based on Spatial-Channel Attention and Semantic Redescription

New Ideas and Trends in Deep Multimodal Content Understanding: A Review

Investigation related to application of Generative Adversarial Networks in text-to-image synthesis

Multimodal Image Synthesis and Editing: The Generative AI Era

Text-to-Image Synthesis With Generative Models: Methods, Datasets, Performance Metrics, Challenges, and Future Direction

RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model

AI-based text-to-image synthesis: A review

A survey on advancements in image-text multimodal models: From general techniques to biomedical implementations

DMF-GAN: Deep Multimodal Fusion Generative Adversarial Networks for Text-to-Image Synthesis