Abstract:An antithetical concept, adaptive symmetry, to conservative symmetry in physics is proposed to understand the deep neural networks (DNNs). It characterizes the invariance of variance, where a biotic system explores different pathways of evolution with equal probability in absence of feedback signals, and complex functional structure emerges from quantitative accumulation of adaptive-symmetries breaking in response to feedback signals. Theoretically and experimentally, we characterize the optimization process of a DNN system as an extended adaptive-symmetry-breaking process. One particular finding is that a hierarchically large DNN would have a large reservoir of adaptive symmetries, and when the information capacity of the reservoir exceeds the complexity of the dataset, the system could absorb all perturbations of the examples and self-organize into a functional structure of zero training errors measured by a certain surrogate risk. More specifically, this process is characterized by a statistical-mechanical model that could be appreciated as a generalization of statistics physics to the DNN organized complex system, and characterizes regularities in higher dimensionality. The model consists of three constitutes that could be appreciated as the counterparts of Boltzmann distribution, Ising model, and conservative symmetry, respectively: (1) a stochastic definition/interpretation of DNNs that is a multilayer probabilistic graphical model, (2) a formalism of circuits that perform biological computation, (3) a circuit symmetry from which self-similarity between the microscopic and the macroscopic adaptability manifests. The model is analyzed with a method referred as the statistical assembly method that analyzes the coarse-grained behaviors (over a symmetry group) of the heterogeneous hierarchical many-body interaction in DNNs.

On the instability and degeneracy of deep learning models

Measuring and Mitigating Local Instability in Deep Neural Networks

How more data can hurt: Instability and regularization in next-generation reservoir computing

Multiplicative noise and heavy tails in stochastic optimization

Instability and Information

Stability Theory of Stochastic Models in Opinion Dynamics

Deep Unsupervised Learning using Nonequilibrium Thermodynamics

Can Stability be Detrimental? Better Generalization through Gradient Descent Instabilities

The Epistemic Uncertainty Hole: an issue of Bayesian Neural Networks

Complexity from Adaptive-Symmetries Breaking: Global Minima in the Statistical Mechanics of Deep Neural Networks

A Survey on Statistical Theory of Deep Learning: Approximation, Training Dynamics, and Generative Models

Data-Dependence of Plateau Phenomenon in Learning with Neural Network --- Statistical Mechanical Analysis

Inconsistency, Instability, and Generalization Gap of Deep Neural Network Training

Deep Learning the Ising Model Near Criticality

Loss as the Inconsistency of a Probabilistic Dependency Graph: Choose Your Model, Not Your Loss Function

Depth Degeneracy in Neural Networks: Vanishing Angles in Fully Connected ReLU Networks on Initialization

Perturbation Analysis of Neural Collapse

A Probabilistic Theory of Deep Learning

Towards Understanding Generalization via Decomposing Excess Risk Dynamics

Prediction Instability in Machine Learning Ensembles

A Theory on Adam Instability in Large-Scale Machine Learning