Abstract:Recently, {\it stochastic momentum} methods have been widely adopted in training deep neural networks. However, their convergence analysis is still underexplored at the moment, in particular for non-convex optimization. This paper fills the gap between practice and theory by developing a basic convergence analysis of two stochastic momentum methods, namely stochastic heavy-ball method and the stochastic variant of Nesterov's accelerated gradient method. We hope that the basic convergence results developed in this paper can serve the reference to the convergence of stochastic momentum methods and also serve the baselines for comparison in future development of stochastic momentum methods. The novelty of convergence analysis presented in this paper is a unified framework, revealing more insights about the similarities and differences between different stochastic momentum methods and stochastic gradient method. The unified framework exhibits a continuous change from the gradient method to Nesterov's accelerated gradient method and finally the heavy-ball method incurred by a free parameter, which can help explain a similar change observed in the testing error convergence behavior for deep learning. Furthermore, our empirical results for optimizing deep neural networks demonstrate that the stochastic variant of Nesterov's accelerated gradient method achieves a good tradeoff (between speed of convergence in training error and robustness of convergence in testing error) among the three stochastic methods.

A Diffusion Approximation Theory of Momentum SGD in Nonconvex Optimization

Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting

Convergence of SGD with momentum in the nonconvex case: A time window-based analysis

On the Convergence of Memory-Based Distributed SGD.

Random Scaling and Momentum for Non-smooth Non-convex Optimization

Stochastic Gradient Descent in the Viewpoint of Graduated Optimization

Error estimates between SGD with momentum and underdamped Langevin diffusion

$μ^2$-SGD: Stable Stochastic Optimization via a Double Momentum Mechanism

Unified Convergence Analysis of Stochastic Momentum Methods for Convex and Non-convex Optimization

A Unified Momentum-based Paradigm of Decentralized SGD for Non-Convex Models and Heterogeneous Data

When and Why Momentum Accelerates SGD:An Empirical Study

Bound Analysis of Natural Gradient Descent in Stochastic Optimization Setting

Shuffling Momentum Gradient Algorithm for Convex Optimization

The Marginal Value of Momentum for Small Learning Rate SGD

A Unified Analysis of Stochastic Momentum Methods for Deep Learning

Learning-rate-free Momentum SGD with Reshuffling Converges in Nonsmooth Nonconvex Optimization

The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization

Stochastic normalized gradient descent with momentum for large-batch training

The Anytime Convergence of Stochastic Gradient Descent with Momentum: From a Continuous-Time Perspective

Acceleration of stochastic gradient descent with momentum by averaging: finite-sample rates and asymptotic normality

Scaling transition from momentum stochastic gradient descent to plain stochastic gradient descent