Kaustubh D. Dhole,Varun Gangal,Sebastian Gehrmann,Aadesh Gupta,Zhenhao Li,Saad Mahamood,Abinaya Mahendiran,Simon Mille,Ashish Shrivastava,Samson Tan,Tongshuang Wu,Jascha Sohl-Dickstein,Jinho D. Choi,Eduard Hovy,Ondrej Dusek,Sebastian Ruder,Sajant Anand,Nagender Aneja,Rabin Banjade,Lisa Barthe,Hanna Behnke,Ian Berlot-Attwell,Connor Boyle,Caroline Brun,Marco Antonio Sobrevilla Cabezudo,Samuel Cahyawijaya,Emile Chapuis,Wanxiang Che,Mukund Choudhary,Christian Clauss,Pierre Colombo,Filip Cornell,Gautier Dagan,Mayukh Das,Tanay Dixit,Thomas Dopierre,Paul-Alexis Dray,Suchitra Dubey,Tatiana Ekeinhor,Marco Di Giovanni,Tanya Goyal,Rishabh Gupta,Rishabh Gupta,Louanes Hamla,Sang Han,Fabrice Harel-Canada,Antoine Honore,Ishan Jindal,Przemyslaw K. Joniak,Denis Kleyko,Venelin Kovatchev,Kalpesh Krishna,Ashutosh Kumar,Stefan Langer,Seungjae Ryan Lee,Corey James Levinson,Hualou Liang,Kaizhao Liang,Zhexiong Liu,Andrey Lukyanenko,Vukosi Marivate,Gerard de Melo,Simon Meoni,Maxime Meyer,Afnan Mir,Nafise Sadat Moosavi,Niklas Muennighoff,Timothy Sum Hon Mun,Kenton Murray,Marcin Namysl,Maria Obedkova,Priti Oli,Nivranshu Pasricha,Jan Pfister,Richard Plant,Vinay Prabhu,Vasile Pais,Libo Qin,Shahab Raji,Pawan Kumar Rajpoot,Vikas Raunak,Roy Rinberg,Nicolas Roberts,Juan Diego Rodriguez,Claude Roux,Vasconcellos P. H. S.,Ananya B. Sai,Robin M. Schmidt,Thomas Scialom,Tshephisho Sefara,Saqib N. Shamsi,Xudong Shen,Haoyue Shi,Yiwen Shi,Anna Shvets,Nick Siegel,Damien Sileo,Jamie Simon,Chandan Singh,Roman Sitelew,Priyank Soni,Taylor Sorensen,William Soto,Aman Srivastava,KV Aditya Srivatsa,Tony Sun,Mukund Varma T,A Tabassum,Fiona Anting Tan,Ryan Teehan,Mo Tiwari,Marie Tolkiehn,Athena Wang,Zijian Wang,Gloria Wang,Zijie J. Wang,Fuxuan Wei,Bryan Wilie,Genta Indra Winata,Xinyi Wu,Witold Wydmański,Tianbao Xie,Usama Yaseen,Michael A. Yee,Jing Zhang,Yue Zhang,et al. (26 additional authors not shown)

Abstract:Data augmentation is an important component in the robustness evaluation of models in natural language processing (NLP) and in enhancing the diversity of the data they are trained on. In this paper, we present NL-Augmenter, a new participatory Python-based natural language augmentation framework which supports the creation of both transformations (modifications to the data) and filters (data splits according to specific features). We describe the framework and an initial set of 117 transformations and 23 filters for a variety of natural language tasks. We demonstrate the efficacy of NL-Augmenter by using several of its transformations to analyze the robustness of popular natural language models. The infrastructure, datacards and robustness analysis results are available publicly on the NL-Augmenter repository (<a class="link-external link-https" href="https://github.com/GEM-benchmark/NL-Augmenter" rel="external noopener nofollow">this https URL</a>).

Evaluation Metrics for Text Data Augmentation in NLP

Task-driven Augmented Data Evaluation

A Survey of Data Augmentation Approaches for NLP

A Survey of Evaluation Metrics Used for NLG Systems

On the Effectiveness of Automated Metrics for Text Generation Systems

Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement Theory

Exploring Data Augmentation Methods on Social Media Corpora

Text Data Augmentation for Deep Learning

Data augmentation in natural language processing: a novel text generation approach for long and short text classifiers

NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference Checklist

A Comprehensive Survey on Data Augmentation

Automatic Metrics in Natural Language Generation: A Survey of Current Evaluation Practices

To Augment or Not to Augment? A Comparative Study on Text Augmentation Techniques for Low-Resource NLP

An Empirical Survey of Data Augmentation for Limited Data Learning in NLP

Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics

On Evaluation Protocols for Data Augmentation in a Limited Data Scenario

Compression, Transduction, and Creation: A Unified Framework for Evaluating Natural Language Generation

NL-Augmenter: A Framework for Task-Sensitive Natural Language Augmentation

BEAMetrics: A Benchmark for Language Generation Evaluation Evaluation

Advancing NLP Models with Strategic Text Augmentation: A Comprehensive Study of Augmentation Methods and Curriculum Strategies

Investigating the Effectiveness of Data Augmentation from Similarity and Diversity: an Empirical Study