CLEAN-EVAL: Clean Evaluation on Contaminated Large Language Models

Wenhong Zhu,Hongkun Hao,Zhiwei He,Yunze Song,Yumeng Zhang,Hanxu Hu,Yiran Wei,Rui Wang,Hongyuan Lu

2024-06-03

Abstract:We are currently in an era of fierce competition among various large language models (LLMs) continuously pushing the boundaries of benchmark performance. However, genuinely assessing the capabilities of these LLMs has become a challenging and critical issue due to potential data contamination, and it wastes dozens of time and effort for researchers and engineers to download and try those contaminated models. To save our precious time, we propose a novel and useful method, Clean-Eval, which mitigates the issue of data contamination and evaluates the LLMs in a cleaner manner. Clean-Eval employs an LLM to paraphrase and back-translate the contaminated data into a candidate set, generating expressions with the same meaning but in different surface forms. A semantic detector is then used to filter the generated low-quality samples to narrow down this candidate set. The best candidate is finally selected from this set based on the BLEURT score. According to human assessment, this best candidate is semantically similar to the original contamination data but expressed differently. All candidates can form a new benchmark to evaluate the model. Our experiments illustrate that Clean-Eval substantially restores the actual evaluation results on contaminated LLMs under both few-shot learning and fine-tuning scenarios.

Computation and Language

What problem does this paper attempt to address?

This paper focuses on the evaluation problem of large-scale language models (LLMs) under training data contamination. Since LLMs are often trained on data extracted from websites and public datasets, there may be overlap between the training data and evaluation benchmarks, resulting in data contamination and overestimation of model performance. The paper proposes a new method called Clean-Eval to address this issue. Clean-Eval utilizes a neural network model to rewrite and translate the contaminated data, generating a candidate set with the same meaning but different expressions. Then, low-quality samples are filtered out by a semantic detector, and the final evaluation set is selected based on BLEURT scores. Experiments show that Clean-Eval effectively restores the actual performance of contaminated models on different tasks, making evaluations more accurate. The paper also emphasizes the significance of clean evaluation in driving the development of the LLMs community.

CLEAN-EVAL: Clean Evaluation on Contaminated Large Language Models

Clean Evaluations on Contaminated Visual Language Models

Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models

Data Contamination Can Cross Language Barriers

An Open Source Data Contamination Report for Large Language Models

KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models

How Much are Large Language Models Contaminated? A Comprehensive Survey and the LLMSanitize Library

CLEVA: Chinese Language Models EVAluation Platform

Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

Concerned with Data Contamination? Assessing Countermeasures in Code Language Model

Investigating Data Contamination in Modern Benchmarks for Large Language Models

Don't Make Your LLM an Evaluation Benchmark Cheater

LLMEval: A Preliminary Study on How to Evaluate Large Language Models

Mitigating the Bias of Large Language Model Evaluation

LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test Construction

Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges

Benchmark Data Contamination of Large Language Models: A Survey

NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

IterClean: an Iterative Data Cleaning Framework with Large Language Models