Abstract:Retrieval-augmented generation (RAG) techniques leverage the in-context learning capabilities of large language models (LLMs) to produce more accurate and relevant responses. Originating from the simple 'retrieve-then-read' approach, the RAG framework has evolved into a highly flexible and modular paradigm. A critical component, the Query Rewriter module, enhances knowledge retrieval by generating a search-friendly query. This method aligns input questions more closely with the knowledge base. Our research identifies opportunities to enhance the Query Rewriter module to Query Rewriter+ by generating multiple queries to overcome the Information Plateaus associated with a single query and by rewriting questions to eliminate Ambiguity, thereby clarifying the underlying intent. We also find that current RAG systems exhibit issues with Irrelevant Knowledge; to overcome this, we propose the Knowledge Filter. These two modules are both based on the instruction-tuned Gemma-2B model, which together enhance response quality. The final identified issue is Redundant Retrieval; we introduce the Memory Knowledge Reservoir and the Retriever Trigger to solve this. The former supports the dynamic expansion of the RAG system's knowledge base in a parameter-free manner, while the latter optimizes the cost for accessing external knowledge, thereby improving resource utilization and response efficiency. These four RAG modules synergistically improve the response quality and efficiency of the RAG system. The effectiveness of these modules has been validated through experiments and ablation studies across six common QA datasets. The source code can be accessed at <a class="link-external link-https" href="https://github.com/Ancientshi/ERM4" rel="external noopener nofollow">this https URL</a>.

Pistis-RAG: A Scalable Cascading Framework Towards Trustworthy Retrieval-Augmented Generation

Pistis-RAG: Enhancing Retrieval-Augmented Generation with Human Feedback

Trustworthiness in Retrieval-Augmented Generation Systems: A Survey

WeKnow-RAG: An Adaptive Approach for Retrieval-Augmented Generation Integrating Web Search and Knowledge Graphs

Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models

RAGLAB: A Modular and Research-Oriented Unified Framework for Retrieval-Augmented Generation

Towards Understanding Retrieval Accuracy and Prompt Quality in RAG Systems

RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation

Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation

RAG Foundry: A Framework for Enhancing LLMs for Retrieval Augmented Generation

LLMs Know What They Need: Leveraging a Missing Information Guided Framework to Empower Retrieval-Augmented Generation

Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting

Enhancing Retrieval and Managing Retrieval: A Four-Module Synergy for Improved Quality and Efficiency in RAG Systems

A Theory for Token-Level Harmonization in Retrieval-Augmented Generation

C-RAG: Certified Generation Risks for Retrieval-Augmented Language Models

RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs

EasyRAG: Efficient Retrieval-Augmented Generation Framework for Automated Network Operations

LightRAG: Simple and Fast Retrieval-Augmented Generation

RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation

DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Models

A Hybrid RAG System with Comprehensive Enhancement on Complex Reasoning