Abstract:The development of Artificial Intelligence Generated Content (AIGC) has been facilitated by advancements in model algorithms, scalable foundation model architectures, and the availability of ample high-quality datasets. While AIGC has achieved remarkable performance, it still faces challenges, such as the difficulty of maintaining up-to-date and long-tail knowledge, the risk of data leakage, and the high costs associated with training and inference. Retrieval-Augmented Generation (RAG) has recently emerged as a paradigm to address such challenges. In particular, RAG introduces the information retrieval process, which enhances AIGC results by retrieving relevant objects from available data stores, leading to greater accuracy and robustness. In this paper, we comprehensively review existing efforts that integrate RAG technique into AIGC scenarios. We first classify RAG foundations according to how the retriever augments the generator. We distill the fundamental abstractions of the augmentation methodologies for various retrievers and generators. This unified perspective encompasses all RAG scenarios, illuminating advancements and pivotal technologies that help with potential future progress. We also summarize additional enhancements methods for RAG, facilitating effective engineering and implementation of RAG systems. Then from another view, we survey on practical applications of RAG across different modalities and tasks, offering valuable references for researchers and practitioners. Furthermore, we introduce the benchmarks for RAG, discuss the limitations of current RAG systems, and suggest potential directions for future research. Project: https://github.com/hymie122/RAG-Survey

Retrieving Multimodal Information for Augmented Generation: A Survey

Retrieving Multimodal Information for Augmented Generation: A Survey

Advanced Embedding Techniques in Multimodal Retrieval Augmented Generation A Comprehensive Study on Cross Modal AI Applications

A Survey of Multimodal Composite Editing and Retrieval

Multimodal Image Synthesis and Editing: The Generative AI Era

LLMs Meet Multimodal Generation and Editing: A Survey

Retrieval-Augmented Generation for AI-Generated Content: A Survey

The 2nd International Workshop on Deep Multi-modal Generation and Retrieval

MORE: Multi-mOdal REtrieval Augmented Generative Commonsense Reasoning

Retrieval-Augmented Generation for Natural Language Processing: A Survey

Retrieval-Augmented Generation for Large Language Models: A Survey

A Survey on Multimodal Benchmarks: In the Era of Large AI Models

Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions

Cross-Modal Knowledge Discovery, Inference, and Challenges.

Multi-modal Retrieval Augmented Multi-modal Generation: A Benchmark, Evaluate Metrics and Strong Baselines

Searching for Best Practices in Retrieval-Augmented Generation

Retrieval-Augmented Multimodal Language Modeling

MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models

Self-adaptive Multimodal Retrieval-Augmented Generation