Abstract:Retrieval-augmented generation (RAG) techniques leverage the in-context learning capabilities of large language models (LLMs) to produce more accurate and relevant responses. Originating from the simple 'retrieve-then-read' approach, the RAG framework has evolved into a highly flexible and modular paradigm. A critical component, the Query Rewriter module, enhances knowledge retrieval by generating a search-friendly query. This method aligns input questions more closely with the knowledge base. Our research identifies opportunities to enhance the Query Rewriter module to Query Rewriter+ by generating multiple queries to overcome the Information Plateaus associated with a single query and by rewriting questions to eliminate Ambiguity, thereby clarifying the underlying intent. We also find that current RAG systems exhibit issues with Irrelevant Knowledge; to overcome this, we propose the Knowledge Filter. These two modules are both based on the instruction-tuned Gemma-2B model, which together enhance response quality. The final identified issue is Redundant Retrieval; we introduce the Memory Knowledge Reservoir and the Retriever Trigger to solve this. The former supports the dynamic expansion of the RAG system's knowledge base in a parameter-free manner, while the latter optimizes the cost for accessing external knowledge, thereby improving resource utilization and response efficiency. These four RAG modules synergistically improve the response quality and efficiency of the RAG system. The effectiveness of these modules has been validated through experiments and ablation studies across six common QA datasets. The source code can be accessed at <a class="link-external link-https" href="https://github.com/Ancientshi/ERM4" rel="external noopener nofollow">this https URL</a>.

EfficientRAG: Efficient Retriever for Multi-Hop Question Answering

Retrieve, Summarize, Plan: Advancing Multi-hop Question Answering with an Iterative Approach

MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries

Hierarchical Retrieval-Augmented Generation Model with Rethink for Multi-hop Question Answering

A Multi-Source Retrieval Question Answering Framework Based on RAG

LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering

EasyRAG: Efficient Retrieval-Augmented Generation Framework for Automated Network Operations

DR-RAG: Applying Dynamic Document Relevance to Retrieval-Augmented Generation for Question-Answering

RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation

MBA-RAG: a Bandit Approach for Adaptive Retrieval-Augmented Generation through Question Complexity

Boosting Conversational Question Answering with Fine-Grained Retrieval-Augmentation and Self-Check

Toward Optimal Search and Retrieval for RAG

Efficient In-Domain Question Answering for Resource-Constrained Environments

W-RAG: Weakly Supervised Dense Retrieval in RAG for Open-domain Question Answering

RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation

Enhancing Retrieval and Managing Retrieval: A Four-Module Synergy for Improved Quality and Efficiency in RAG Systems

Optimizing Retrieval-Augmented Generation with Elasticsearch for Enhanced Question-Answering Systems

Multi-Meta-RAG: Improving RAG for Multi-Hop Queries using Database Filtering with LLM-Extracted Metadata

IM-RAG: Multi-Round Retrieval-Augmented Generation Through Learning Inner Monologues

ERAGent: Enhancing Retrieval-Augmented Language Models with Improved Accuracy, Efficiency, and Personalization

Retrieval-Augmented Generation for Domain-Specific Question Answering: A Case Study on Pittsburgh and CMU