Gradient Checkpointing

Gradient checkpointing is a memory-saving technique used when training deep learning models. Instead of saving all intermediate layer activations during the forward pass, it discards most of them and recomputes them on-the-fly during the backward pass. This trades a small increase in computation time for a massive reduction in VRAM usage

How Gradient Checkpointing Works

Often combined with Mixed Precision to further cut training memory.


References