Metadata-based Data Exploration with Retrieval-Augmented Generation for Large Language Models

Teruaki Hayashi,Hiroki Sakaji,Jiayi Dai,Randy Goebel

2024-10-06

Abstract:Developing the capacity to effectively search for requisite datasets is an urgent requirement to assist data users in identifying relevant datasets considering the very limited available metadata. For this challenge, the utilization of third-party data is emerging as a valuable source for improvement. Our research introduces a new architecture for data exploration which employs a form of Retrieval-Augmented Generation (RAG) to enhance metadata-based data discovery. The system integrates large language models (LLMs) with external vector databases to identify semantic relationships among diverse types of datasets. The proposed framework offers a new method for evaluating semantic similarity among heterogeneous data sources and for improving data exploration. Our study includes experimental results on four critical tasks: 1) recommending similar datasets, 2) suggesting combinable datasets, 3) estimating tags, and 4) predicting variables. Our results demonstrate that RAG can enhance the selection of relevant datasets, particularly from different categories, when compared to conventional metadata approaches. However, performance varied across tasks and models, which confirms the significance of selecting appropriate techniques based on specific use cases. The findings suggest that this approach holds promise for addressing challenges in data exploration and discovery, although further refinement is necessary for estimation tasks.

Information Retrieval

What problem does this paper attempt to address?

The paper attempts to address the problem of how to effectively search and identify relevant datasets when existing metadata is limited. Specifically, the study proposes a new data exploration framework that leverages Retrieval-Augmented Generation (RAG) technology to enhance metadata-based data discovery capabilities. This approach aims to overcome the limitations of traditional metadata methods in data search, particularly for users lacking specific domain knowledge, where conventional query matching methods often struggle to find the most relevant datasets. Additionally, the study explores how to evaluate the semantic similarity between heterogeneous data and experimentally validates the effectiveness of RAG in tasks such as recommending similar datasets, suggesting composable datasets, estimating labels, and predicting variables. The research results indicate that although the performance varies across different tasks and models, RAG shows potential in selecting relevant datasets across categories. Therefore, the core question of the paper is: How to evaluate the semantic similarity between heterogeneous data and improve the accuracy of data discovery in the absence of sufficient metadata?

Metadata-based Data Exploration with Retrieval-Augmented Generation for Large Language Models

Retrieval-Augmented Generation for Large Language Models: A Survey

Meta Knowledge for Retrieval Augmented Large Language Models

Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely

Towards Optimizing a Retrieval Augmented Generation using Large Language Model on Academic Data

A Survey on Retrieval-Augmented Text Generation for Large Language Models

Wiping out the limitations of Large Language Models -- A Taxonomy for Retrieval Augmented Generation

Enhancing Retrieval Processes for Language Generation with Augmented Queries

A Method for Parsing and Vectorization of Semi-structured Data used in Retrieval Augmented Generation

Deploying Large Language Models With Retrieval Augmented Generation

Retrieval-Augmented Generation for Natural Language Processing: A Survey

A Survey on RAG Meeting LLMs: Towards Retrieval-Augmented Large Language Models

Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems

Retrieval-Generation Synergy Augmented Large Language Models

LightRAG: Simple and Fast Retrieval-Augmented Generation

Making Metadata More FAIR Using Large Language Models

Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts

Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report

Enhancing Retrieval-Augmented Large Language Models with Iterative Retrieval-Generation Synergy

Auto-RAG: Autonomous Retrieval-Augmented Generation for Large Language Models