Abstract:Training in Feed Forward Deep Neural Networks is a memory-intensive operation which is usually performed on GPUs with limited memory capacities. This may force data scientists to limit the depth of the models or the resolution of the input data if data does not fit in the GPU memory. The re-materialization technique, whose idea comes from the checkpointing strategies developed in the Automatic Differentiation literature, allows data scientists to limit the memory requirements related to the storage of intermediate data (activations), at the cost of an increase in the computational cost. This paper introduces a new strategy of re-materialization of activations that significantly reduces memory usage. It consists in selecting which activations are saved and which activations are deleted during the forward phase, and then recomputing the deleted activations when they are needed during the backward phase. We propose an original computation model that combines two types of activation savings: either only storing the layer inputs, or recording the complete history of operations that produced the outputs. This paper focuses on the fully heterogeneous case, where the computation time and the memory requirement of each layer is different. We prove that finding the optimal solution is NP-hard and that classical techniques from Automatic Differentiation literature do not apply. Moreover, the classical assumption of memory persistence of materialized activations, used to simplify the search of optimal solutions, does not hold anymore. Thus, we propose a weak memory persistence property and provide a Dynamic Program to compute the optimal sequence of computations. This algorithm is made available through the Rotor software, a PyTorch plug-in dealing with any network consisting of a sequence of layers, each of them having an arbitrarily complex structure. Through extensive experiments, we show that our implementation consistently outperforms existing re-materialization approaches for a large class of networks, image sizes and batch sizes.

STR: Hybrid Tensor Re-Generation to Break Memory Wall for DNN Training

A Swap Dominated Tensor Re-Generation Strategy for Training Deep Learning Models

TSPLIT: Fine-grained GPU Memory Management for Efficient DNN Training Via Tensor Splitting

MegTaiChi: Dynamic Tensor-based Memory Management Optimization for DNN Training

Efficient Memory Management for GPU-based Deep Learning Systems

MAGIS: Memory Optimization Via Coordinated Graph Transformation and Scheduling for DNN

pommDNN: Performance optimal GPU memory management for deep neural network training

Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers

SuperNeurons: Dynamic GPU Memory Management for Training Deep Neural Networks

Optimization of GPU Memory Usage for Training Deep Neural Networks.

Optimal Re-Materialization Strategies for Heterogeneous Chains: How to Train Deep Neural Networks with Limited Memory

Accelerating Tensor Swapping in GPUs with Self-Tuning Compression

Memory Optimization for Deep Networks

Coop: Memory is not a Commodity

Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading

G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor Migrations

HOME: A Holistic GPU Memory Management Framework for Deep Learning

ROAM: memory-efficient large DNN training via optimized operator ordering and memory layout

Improving Automatic Parallel Training Via Balanced Memory Workload Optimization

Hybrid Tensor Decomposition in Neural Network Compression

WELDER: Scheduling Deep Learning Memory Access Via Tile-graph