Abstract:Training in Feed Forward Deep Neural Networks is a memory-intensive operation which is usually performed on GPUs with limited memory capacities. This may force data scientists to limit the depth of the models or the resolution of the input data if data does not fit in the GPU memory. The re-materialization technique, whose idea comes from the checkpointing strategies developed in the Automatic Differentiation literature, allows data scientists to limit the memory requirements related to the storage of intermediate data (activations), at the cost of an increase in the computational cost. This paper introduces a new strategy of re-materialization of activations that significantly reduces memory usage. It consists in selecting which activations are saved and which activations are deleted during the forward phase, and then recomputing the deleted activations when they are needed during the backward phase. We propose an original computation model that combines two types of activation savings: either only storing the layer inputs, or recording the complete history of operations that produced the outputs. This paper focuses on the fully heterogeneous case, where the computation time and the memory requirement of each layer is different. We prove that finding the optimal solution is NP-hard and that classical techniques from Automatic Differentiation literature do not apply. Moreover, the classical assumption of memory persistence of materialized activations, used to simplify the search of optimal solutions, does not hold anymore. Thus, we propose a weak memory persistence property and provide a Dynamic Program to compute the optimal sequence of computations. This algorithm is made available through the Rotor software, a PyTorch plug-in dealing with any network consisting of a sequence of layers, each of them having an arbitrarily complex structure. Through extensive experiments, we show that our implementation consistently outperforms existing re-materialization approaches for a large class of networks, image sizes and batch sizes.

Adaptive higher order reversible integrators for memory efficient deep learning

Neural Operator Learning for Long-Time Integration in Dynamical Systems with Recurrent Neural Networks

Efficient, Accurate and Stable Gradients for Neural ODEs

High-Performance Temporal Reversible Spiking Neural Networks with $O(L)$ Training Memory and $O(1)$ Inference Cost

Time Dependence in Non-Autonomous Neural ODEs

Do Residual Neural Networks discretize Neural Ordinary Differential Equations?

Training Stiff Neural Ordinary Differential Equations with Explicit Exponential Integration Methods

Optimal Re-Materialization Strategies for Heterogeneous Chains: How to Train Deep Neural Networks with Limited Memory

Reversible Recurrent Neural Networks

Understanding Latent Timescales in Neural Ordinary Differential Equation Models for Advection-Dominated Dynamical Systems

A memory-efficient neural ODE framework based on high-level adjoint differentiation

Memory-Efficient Reversible Spiking Neural Networks

Deep Recurrent Neural Network Architecture of High Order Indirect Integration Method

A Forward Learning Algorithm for Neural Memory Ordinary Differential Equations

Systematic construction of continuous-time neural networks for linear dynamical systems

Interpretable learning of effective dynamics for multiscale systems

A Stable and Scalable Method for Solving Initial Value PDEs with Neural Networks

Neural Integration of Continuous Dynamics

Revising the Structure of Recurrent Neural Networks to Eliminate Numerical Derivatives in Forming Physics Informed Loss Terms with Respect to Time

Neural Ordinary Differential Equations for Model Order Reduction of Stiff Systems

Dynamical System Inspired Adaptive Time Stepping Controller for Residual Network Families