Abstract:We propose a novel automatic testing method, hybrid mutation driven testing (HMT), which extends the mutation idea in natural language inference (NLI). We apply four mutation operators to achieve the hybrid mutation strategy, mutating the premise and the hypothesis in the samples jointly or individually. The experimental results show that HMT can effectively generate mutations and trigger the inconsistency bugs of NLI models, with independent bugs for four mutation operators. Summary Natural language inference (NLI) is a task to infer the relationship between the premise and hypothesis sentences, whose models have essential applications in the many natural language processing (NLP) fields, for example, machine reading comprehension and recognizing textual entailment. Due to the data‐driven programming paradigm, bugs inevitably occur in NLI models during the application process, which calls for novel automatic testing techniques to deal with NLI testing challenges. The main difficulty in achieving automatic testing for NLI models is the oracle problem; that is, it may be too expensive to label NLI model inputs manually and hence be too challenging to verify the correctness of model outputs. To tackle the oracle problem, this study proposes a novel automatic testing method hybrid mutation driven testing (HMT), which extends the mutation idea applied in other NLP domains successfully. Specifically, as there are two sets of sentences, that is, premise and hypothesis, to be mutated, we propose four mutation operators to achieve the hybrid mutation strategy, which mutate the premise and the hypothesis sentences jointly or individually. We assume that the mutation would not affect the outputs; that is, if the original and mutated outputs are inconsistent, inconsistency bugs could be detected without knowing the true labels. To evaluate our method HMT, we conduct experiments on two widely used datasets with two advanced models and generate more than 520,000 mutations by applying our mutation operators. Our experimental results show that (a) our method, HMT, can effectively generate mutated testing samples, (b) our method can effectively trigger the inconsistency bugs of the NLI models, and (c) all four mutation operators can independently trigger inconsistency bugs.

Natural Test Generation for Precise Testing of Question Answering Software.

QATest: A Uniform Fuzzing Framework for Question Answering Systems.

Knowledge Graph Driven Inference Testing for Question Answering Software

Quality Assurance of Bioinformatics Software: A Case Study of Testing a Biomedical Text Processing Tool Using Metamorphic Testing

Syn-QA2: Evaluating False Assumptions in Long-tail Questions with Synthetic QA Datasets

Supporting maintenance and testing for AI functions of mobile apps based on user reviews: An empirical study on plant identification apps

(QA)$^2$: Question Answering with Questionable Assumptions

Can ChatGPT advance software testing intelligence? An experience report on metamorphic testing

QADYNAMICS: Training Dynamics-Driven Synthetic QA Diagnostic for Zero-Shot Commonsense Question Answering

Automatic Test Generation for Mutation Testing on Database Applications.

Analysis of QA System Behavior against Context and Question Changes

AI-based Question Answering Assistance for Analyzing Natural-language Requirements

Automatic Question Generation for Repeated Testing to Improve Student Learning Outcome

On the Generation of Medical Question-Answer Pairs

Improving the Quality of Computational Science Software by Using Metamorphic Relations to Test Machine Learning Applications

Intergenerational Test Generation for Natural Language Processing Applications

Hybrid mutation driven testing for natural language inference

Structure and Performance Analysis of Open Domain QA System

KBQA: Learning Question Answering over QA Corpora and Knowledge Bases

Aggregated Knowledge Model: Enhancing Domain-Specific QA with Fine-Tuned and Retrieval-Augmented Generation Models