Abstract:We propose a novel automatic testing method, hybrid mutation driven testing (HMT), which extends the mutation idea in natural language inference (NLI). We apply four mutation operators to achieve the hybrid mutation strategy, mutating the premise and the hypothesis in the samples jointly or individually. The experimental results show that HMT can effectively generate mutations and trigger the inconsistency bugs of NLI models, with independent bugs for four mutation operators. Summary Natural language inference (NLI) is a task to infer the relationship between the premise and hypothesis sentences, whose models have essential applications in the many natural language processing (NLP) fields, for example, machine reading comprehension and recognizing textual entailment. Due to the data‐driven programming paradigm, bugs inevitably occur in NLI models during the application process, which calls for novel automatic testing techniques to deal with NLI testing challenges. The main difficulty in achieving automatic testing for NLI models is the oracle problem; that is, it may be too expensive to label NLI model inputs manually and hence be too challenging to verify the correctness of model outputs. To tackle the oracle problem, this study proposes a novel automatic testing method hybrid mutation driven testing (HMT), which extends the mutation idea applied in other NLP domains successfully. Specifically, as there are two sets of sentences, that is, premise and hypothesis, to be mutated, we propose four mutation operators to achieve the hybrid mutation strategy, which mutate the premise and the hypothesis sentences jointly or individually. We assume that the mutation would not affect the outputs; that is, if the original and mutated outputs are inconsistent, inconsistency bugs could be detected without knowing the true labels. To evaluate our method HMT, we conduct experiments on two widely used datasets with two advanced models and generate more than 520,000 mutations by applying our mutation operators. Our experimental results show that (a) our method, HMT, can effectively generate mutated testing samples, (b) our method can effectively trigger the inconsistency bugs of the NLI models, and (c) all four mutation operators can independently trigger inconsistency bugs.

Hybrid mutation driven testing for natural language inference

MUT: Human-in-the-Loop Unit Test Migration

Training NLI Models Through Universal Adversarial Attack

Learning Likely Invariants to Explain Why a Program Fails

An Exploratory Study on Using Large Language Models for Mutation Testing

MILE: A Mutation Testing Framework of In-Context Learning Systems

Effective test generation using pre-trained Large Language Models and mutation testing

Natural Language Inference Using Lstm Model With Sentence Fusion

Test Case Level Predictive Mutation Testing Combining PIE and Natural Language Features

Intergenerational Test Generation for Natural Language Processing Applications

LLMorpheus: Mutation Testing using Large Language Models

What Are We Really Testing in Mutation Testing for Machine Learning? A Critical Reflection

Natural Test Generation for Precise Testing of Question Answering Software.

Testing Untestable Neural Machine Translation: an Industrial Case

Efficient Mutation Testing via Pre-Trained Language Models

Enhancing Fault Detection for Large Language Models via Mutation-Based Confidence Smoothing

Syntactic Vs. Semantic Similarity of Artificial and Real Faults in Mutation Testing Studies

Evaluation and Improvement of Fault Detection for Large Language Models

Mutation Testing of Unsupervised Learning Systems

DeepMutation: Mutation Testing of Deep Learning Systems