Gradient Checkpointing
Gradient checkpointing is a memory-saving technique used when training deep learning models. Instead of saving all intermediate layer activations during the forward pass, it discards most of them and recomputes them on-the-fly during the backward pass. This trades a small increase in computation time for a massive reduction in VRAM usage
How Gradient Checkpointing Works
-
Standard Forward Pass: Normally, every layer's output (activation) is kept in memory because Back Propagation needs them to calculate gradients.
-
The Checkpointing Trade-off: The model saves only a sparse set of activations ("checkpoints") and deletes the rest to save space.
-
On-the-Fly Recomputation: When the backward pass reaches a missing activation, it re-runs a quick forward pass from the nearest saved checkpoint to recreate the exact values needed.
Often combined with Mixed Precision to further cut training memory.