Abstract:Stochastic gradient descent is one of the most common iterative algorithms used in machine learning and its convergence analysis is a rich area of research. Understanding its convergence properties can help inform what modifications of it to use in different settings. However, most theoretical results either assume convexity or only provide convergence results in mean. This paper, on the other hand, proves convergence bounds in high probability without assuming convexity. Assuming strong smoothness, we prove high probability convergence bounds in two settings: (1) assuming the Polyak-Łojasiewicz inequality and norm sub-Gaussian gradient noise and (2) assuming norm sub-Weibull gradient noise. In the second setting, as an intermediate step to proving convergence, we prove a sub-Weibull martingale difference sequence self-normalized concentration inequality of independent interest. It extends Freedman-type concentration beyond the sub-exponential threshold to heavier-tailed martingale difference sequences. We also provide a post-processing method that picks a single iterate with a provable convergence guarantee as opposed to the usual bound for the unknown best iterate. Our convergence result for sub-Weibull noise extends the regime where stochastic gradient descent has equal or better convergence guarantees than stochastic gradient descent with modifications such as clipping, momentum, and normalization.

Convergence analysis of stochastic higher-order majorization-minimization algorithms

Incremental Majorization-Minimization Optimization with Application to Large-Scale Machine Learning

Stochastic Variance-Reduced Majorization-Minimization Algorithms

On the Global Convergence of Majorization Minimization Algorithms for Nonconvex Optimization Problems.

O(log T) Projections for Stochastic Optimization of Smooth and Strongly Convex Functions

High Probability Convergence Bounds for Non-convex Stochastic Gradient Descent with Sub-Weibull Noise

Can Decentralized Stochastic Minimax Optimization Algorithms Converge Linearly for Finite-Sum Nonconvex-Nonconcave Problems?

On the Local Convergence of a Stochastic Semismooth Newton Method for Nonsmooth Nonconvex Optimization

Convergence in High Probability of Distributed Stochastic Gradient Descent Algorithms

High-Probability Bounds for Stochastic Optimization and Variational Inequalities: the Case of Unbounded Variance

A stochastic use of the Kurdyka-Lojasiewicz property: Investigation of optimization algorithms behaviours in a non-convex differentiable framework

New nonasymptotic convergence rates of stochastic proximal pointalgorithm for convex optimization problems

Convergence of alternating direction method for minimizing sum of two nonconvex functions with linear constraints

On the asymptotic rate of convergence of Stochastic Newton algorithms and their Weighted Averaged versions

Stochastic Optimization for Non-convex Inf-Projection Problems

High-Probability Convergence for Composite and Distributed Stochastic Minimization and Variational Inequalities with Heavy-Tailed Noise

Global Convergence of Langevin Dynamics Based Algorithms for Nonconvex Optimization

On Inhomogeneous Infinite Products of Stochastic Matrices and Applications

Asymptotic Optimality in Stochastic Optimization

Convergence Analysis of Generalized ADMM with Majorization for Linearly Constrained Composite Convex Optimization

Generalized Majorization-Minimization for Non-Convex Optimization.