Mixed Precision Training
A training method combining half-precision (FP16/BF16) and single-precision (FP32) to speed up GPU operations.
Last reviewed: July 25, 2026
Mixed precision training combines lower-precision numeric formats (like FP16 or BF16) with full 32-bit precision (FP32) selectively within a single training run, using each where it’s most beneficial rather than committing entirely to one format. It’s the standard approach for training modern large language models, since it captures most of the speed and memory benefits of low-precision arithmetic while avoiding the numerical instability that pure low-precision training can cause.
How It Works
In a typical mixed precision setup, the bulk of computationally expensive operations — matrix multiplications in attention and feed-forward layers — run in BF16 or FP16, taking advantage of modern GPU tensor cores that execute these operations significantly faster in lower precision. A small number of numerically sensitive operations, however, are kept in FP32: this often includes maintaining a master copy of the model’s weights in FP32 (updated by the optimizer, then cast down to BF16/FP16 for the next forward pass) and certain reduction operations like normalization layers, where accumulating many small values in low precision can lose meaningful accuracy.
Why It’s Necessary
Training purely in FP16 without any FP32 components is prone to numerical problems: gradients can underflow to zero in FP16’s narrower representable range, silently halting learning for affected parameters. Mixed precision training with an FP32 master weight copy (and, for FP16 specifically, loss scaling to keep gradient values within FP16’s representable range) sidesteps this while still executing the expensive bulk of computation at lower precision.
Practical Impact
Mixed precision training typically delivers a 2-3x speedup and roughly halves memory usage for activations compared to pure FP32 training, which is why it’s been standard practice for training large neural networks since GPUs began including dedicated low-precision tensor core hardware.
Automatic Mixed Precision
Most deep learning frameworks provide automatic mixed precision (AMP) tooling that handles the details of mixed precision training with minimal manual intervention — PyTorch’s torch.cuda.amp, for example, automatically decides which operations run in reduced precision and which stay in FP32, and manages loss scaling automatically when using FP16. This has made mixed precision training accessible without requiring practitioners to manually reason about which specific operations are numerically sensitive, reducing what was once a meaningful engineering challenge to often just a few lines of configuration in a standard training script.
This automation is part of why mixed precision became close to a default setting in modern training pipelines within just a few years of its introduction, rather than remaining a specialized optimization only the largest, most sophisticated training teams bothered to implement.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.