Abstract:Recently, {\it stochastic momentum} methods have been widely adopted in training deep neural networks. However, their convergence analysis is still underexplored at the moment, in particular for non-convex optimization. This paper fills the gap between practice and theory by developing a basic convergence analysis of two stochastic momentum methods, namely stochastic heavy-ball method and the stochastic variant of Nesterov's accelerated gradient method. We hope that the basic convergence results developed in this paper can serve the reference to the convergence of stochastic momentum methods and also serve the baselines for comparison in future development of stochastic momentum methods. The novelty of convergence analysis presented in this paper is a unified framework, revealing more insights about the similarities and differences between different stochastic momentum methods and stochastic gradient method. The unified framework exhibits a continuous change from the gradient method to Nesterov's accelerated gradient method and finally the heavy-ball method incurred by a free parameter, which can help explain a similar change observed in the testing error convergence behavior for deep learning. Furthermore, our empirical results for optimizing deep neural networks demonstrate that the stochastic variant of Nesterov's accelerated gradient method achieves a good tradeoff (between speed of convergence in training error and robustness of convergence in testing error) among the three stochastic methods.

Multi-stage stochastic gradient method with momentum acceleration

Stochastic Momentum Method with Double Acceleration for Regularized Empirical Risk Minimization

Stagewise Accelerated Stochastic Gradient Methods for Nonconvex Optimization

Combining Conjugate Gradient and Momentum for Unconstrained Stochastic Optimization With Applications to Machine Learning

Stochastic Gradient Descent with Nonlinear Conjugate Gradient-Style Adaptive Momentum

Optimal Adaptive and Accelerated Stochastic Gradient Descent

An Accelerated Distributed Stochastic Gradient Method with Momentum

Accelerating Asynchronous Algorithms for Convex Optimization by Momentum Compensation

A new non-adaptive optimization method: Stochastic gradient descent with momentum and difference

Delayed supermartingale convergence lemmas for stochastic approximation with Nesterov momentum

Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence

Momentum Schemes with Stochastic Variance Reduction for Nonconvex Composite Optimization

Convergence and Stability of the Stochastic Proximal Point Algorithm with Momentum

Fast Multiobjective Gradient Methods with Nesterov Acceleration via Inertial Gradient-Like Systems

Losing momentum in continuous-time stochastic optimisation

CoolMomentum: A Method for Stochastic Optimization by Langevin Dynamics with Simulated Annealing

Unified Convergence Analysis of Stochastic Momentum Methods for Convex and Non-convex Optimization

Fast Stochastic Variance Reduced Gradient Method with Momentum Acceleration for Machine Learning

Accelerated Stochastic Min-Max Optimization Based on Bias-corrected Momentum

Beneficial effect of anti-interleukin-4 antibody when administered in a murine model of tuberculosis infection.

Accelerated Doubly Stochastic Gradient Algorithm for Large-scale Empirical Risk Minimization