Abstract:Data augmentation is a dominant method for reducing model overfitting and improving generalization. Most existing data augmentation methods tend to find a compromise in augmenting the data, \textit{i.e.}, increasing the amplitude of augmentation carefully to avoid degrading some data too much and doing harm to the model performance. We delve into the relationship between data augmentation and model performance, revealing that the performance drop with heavy augmentation comes from the presence of out-of-distribution (OOD) data. Nonetheless, as the same data transformation has different effects for different training samples, even for heavy augmentation, there remains part of in-distribution data which is beneficial to model training. Based on the observation, we propose a novel data augmentation method, named \textbf{DualAug}, to keep the augmentation in distribution as much as possible at a reasonable time and computational cost. We design a data mixing strategy to fuse augmented data from both the basic- and the heavy-augmentation branches. Extensive experiments on supervised image classification benchmarks show that DualAug improve various automated data augmentation method. Moreover, the experiments on semi-supervised learning and contrastive self-supervised learning demonstrate that our DualAug can also improve related method. Code is available at \href{<a class="link-external link-https" href="https://github.com/shuguang99/DualAug" rel="external noopener nofollow">this https URL</a>}{<a class="link-external link-https" href="https://github.com/shuguang99/DualAug" rel="external noopener nofollow">this https URL</a>}.

Generalization Gap in Data Augmentation: Insights from Illumination

Boosting Unsupervised Contrastive Learning Using Diffusion-Based Data Augmentation from Scratch

Data Augmentation Revisited: Rethinking the Distribution Gap between Clean and Augmented Data

A Simple Feature Augmentation for Domain Generalization

Exploring Data Augmentations on Self-/Semi-/Fully- Supervised Pre-trained Models

On the Generalization Effects of Linear Transformations in Data Augmentation

Effective Data Augmentation With Diffusion Models

Untapped Potential of Data Augmentation: A Domain Generalization Viewpoint

Improving Generalization in Game Agents with Data Augmentation in Imitation Learning

Illumination-Based Data Augmentation for Robust Background Subtraction

Understanding Data Augmentation from a Robustness Perspective

A Preliminary Study on Data Augmentation of Deep Learning for Image Classification

DualAug: Exploiting Additional Heavy Augmentation with OOD Data Rejection

A Good Data Augmentation Policy Is Not All You Need: A Multi-Task Learning Perspective

Image Data Augmentation for Deep Learning: A Survey

Data Augmentation with Illumination Correction in Sematic Segmentation

Generalization to translation shifts: a study in architectures and augmentations

A survey on Image Data Augmentation for Deep Learning

Feature Augmentation for Self-supervised Contrastive Learning: A Closer Look

The Performance Research of the Data Augmentation Method for Image Classification