Abstract:Convolutional Neural networks (CNNs) based applications have become ubiquitous, where proper regularization is greatly needed. To prevent large neural network models from overfitting, dropout has been widely used as an efficient regularization technique in practice. However, many recent works show that the standard dropout is ineffective or even detrimental to the training of CNNs. In this paper, we revisit this issue and examine various dropout variants in an attempt to improve existing dropout-based regularization techniques for CNNs. We attribute the failure of standard dropout to the conflict between the stochasticity of dropout and its following Batch Normalization (BN), and propose to reduce the conflict by placing dropout operations right before the convolutional operation instead of BN, or totally address this issue by replacing BN with Group Normalization (GN). We further introduce a structurally more suited dropout variant Drop-Conv2d, which provides more efficient and effective regularization for deep CNNs. These dropout variants can be readily integrated into the building blocks of CNNs and implemented in existing deep learning platforms. Extensive experiments on benchmark datasets including CIFAR, SVHN and ImageNet are conducted to compare the existing building blocks and the proposed ones with dropout training. Results show that our building blocks improve over state-of-the-art CNNs significantly, which is mainly due to the better regularization and implicit model ensemble effect.

Implicit Regularization of Dropout

Dropout in Training Neural Networks: Flatness of Solution and Noise Structure

Wordreg: Mitigating the Gap Between Training and Inference with Worst-Case Drop Regularization

Stochastic Modified Equations and Dynamics of Dropout Algorithm

A variance principle explains why dropout finds flatter minima

Dropout Training, Data-dependent Regularization, and Generalization Bounds.

Dropout Reduces Underfitting

Rethinking the Usage of Batch Normalization and Dropout in the Training of Deep Neural Networks

R-Drop: Regularized Dropout for Neural Networks.

Asymptotic Convergence Rate of Dropout on Shallow Linear Neural Networks

Dropout as an Implicit Gating Mechanism For Continual Learning

Dropout Rademacher Complexity of Deep Neural Networks.

Y-Drop: A Conductance based Dropout for fully connected layers

Effective and Efficient Dropout for Deep Convolutional Neural Networks

Implicit Regularization in Deep Learning

Implicit Regularization in Deep Matrix Factorization

Dropout, a basic and effective regularization method for a deep learning model: a case study

Dropout: Explicit Forms and Capacity Control

A global convergence theory for deep ReLU implicit networks via over-parameterization

Implicit Regularization in ReLU Networks with the Square Loss

Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning