Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

Activation Checkpointing

A training memory optimization that discards intermediate layer activations, recalculating them during the backward pass.

Last reviewed: July 25, 2026

Technical Overview of Activation Checkpointing

Activation Checkpointing (also known as Gradient Checkpointing) is a memory optimization technique that reduces the GPU memory footprint during neural network training.

During the forward pass of standard training, the GPU must cache intermediate layer activations to compute gradients during the backward pass. For deep models, storing these activations consumes significant VRAM. Activation Checkpointing addresses this by saving only a subset of activations (checkpoints) and recalculating the discarded activations on-demand during the backward pass.

Key Architecture & Implementation

The algorithm trades computational overhead for memory savings:

  1. Forward Pass:
    • The model processes layers, but only saves activation tensors at designated checkpoint layers (e.g., at the boundary of every 4th transformer block).
    • Intermediate activations between checkpoints are discarded.
  2. Backward Pass:
    • When backpropagation reaches a segment with missing activations, the model runs a local forward pass starting from the nearest upstream checkpoint to recalculate the needed tensors.
    • Once gradients are calculated, the recalculated activations are discarded again.
Standard:     F_L1 -> F_L2 -> F_L3 -> F_L4 (All activations saved in VRAM)
Checkpoint:   F_L1 (Save) -> F_L2 (Discard) -> F_L3 (Discard) -> F_L4 (Discard)
              Backward: Recompute L2 & L3 using saved L1 activation

Core Parameters

  • Memory Savings: Can reduce activation memory footprints by up to 60-70%.
  • Compute Penalty: Increases the execution time of the forward pass by approximately 30% due to the recalculation steps.

Real-world Applications

  • Essential for training models with long sequence lengths or large batch sizes.
  • Supported natively in PyTorch and distributed training libraries like DeepSpeed.

The Memory-Compute Tradeoff

During a standard forward and backward pass, a neural network keeps every intermediate layer’s activations in memory throughout the forward pass, because backpropagation needs them to compute gradients on the way back. For very deep networks or long sequences, storing every single intermediate activation can consume more memory than storing the model’s own weights. Activation checkpointing addresses this by deliberately discarding most intermediate activations right after they’re used in the forward pass, keeping only a sparse set of “checkpoints” at select points in the network. During the backward pass, whenever a discarded activation is needed to compute a gradient, the network simply recomputes it on the fly from the nearest saved checkpoint, rather than having stored it the whole time.

Why It’s Worth the Recomputation Cost

This trades a meaningful reduction in memory usage — often 60-80% less activation memory — for extra compute, since some activations end up being computed twice (once during the original forward pass, once during recomputation in the backward pass). In practice, this tradeoff is almost always worthwhile for large model training, because GPU memory is typically the tighter constraint than raw compute time, and the technique is standard in virtually every large-scale training framework, often applied selectively to only the most memory-expensive layers rather than the entire network.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.