Abstract:We develop a geometric approximation theory for deep feed-forward neural networks with ReLU activations. Given a $d$-dimensional hypersurface in $\mathbb{R}^{d+1}$ represented as the graph of a $C^2$-function $\phi$, we show that a deep fully-connected ReLU network of width $d+1$ can implicitly construct an approximation as its zero contour with a precision bound depending on the number of layers. This result is directly applicable to the binary classification setting where the sign of the network is trained as a classifier, with the network's zero contour as a decision boundary. Our proof is constructive and relies on the geometrical structure of ReLU layers provided in [doi:<a class="link-https link-external" data-doi="10.48550/arXiv.2310.03482" href="https://doi.org/10.48550/arXiv.2310.03482" rel="external noopener nofollow">https://doi.org/10.48550/arXiv.2310.03482</a>]. Inspired by this geometrical description, we define a new equivalent network architecture that is easier to interpret geometrically, where the action of each hidden layer is a projection onto a polyhedral cone derived from the layer's parameters. By repeatedly adding such layers, with parameters chosen such that we project small parts of the graph of $\phi$ from the outside in, we, in a controlled way, construct a network that implicitly approximates the graph over a ball of radius $R$. The accuracy of this construction is controlled by a discretization parameter $\delta$ and we show that the tolerance in the resulting error bound scales as $(d-1)R^{3/2}\delta^{1/2}$ and the required number of layers is of order $d\big(\frac{32R}{\delta}\big)^{\frac{d+1}{2}}$.

Universal Function Approximation by Deep Neural Nets with Bounded Width and ReLU Activations

Minimum width for universal approximation using ReLU networks on compact domain

Neural networks with ReLU powers need less depth

New advances in universal approximation with neural networks of minimal width

The Expressive Power of Neural Networks: A View from the Width

Towards Lower Bounds on the Depth of ReLU Neural Networks

Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal Approximation

Universal approximation with complex-valued deep narrow neural networks

On Minimal Depth in Neural Networks

On the Expressive Power of Neural Networks

Optimal Neural Network Approximation for High-Dimensional Continuous Functions

Approximation Error and Complexity Bounds for ReLU Networks on Low-Regular Function Spaces

Implicit Hypersurface Approximation Capacity in Deep ReLU Networks

Optimal Approximation Rates for Deep ReLU Neural Networks on Sobolev and Besov Spaces

Rates of Approximation by ReLU Shallow Neural Networks

Low dimensional approximation and generalization of multivariate functions on smooth manifolds using deep ReLU neural networks

Deep Network Approximation: Achieving Arbitrary Accuracy with Fixed Number of Neurons

On the optimal approximation of Sobolev and Besov functions using deep ReLU neural networks

Deep Neural Networks with ReLU-Sine-Exponential Activations Break Curse of Dimensionality in Approximation on Hölder Class.

Minimal Width for Universal Property of Deep RNN

Why Deep Neural Networks for Function Approximation?