Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

BF16 Precision

A 16-bit floating-point format that matches the dynamic range of FP32, preventing underflow issues.

Last reviewed: July 25, 2026

BF16, or Brain Floating Point 16, is a 16-bit numeric format developed at Google specifically for machine learning workloads, and it has become the default precision for training most modern large language models on GPUs and TPUs. Like standard FP32, BF16 allocates 8 bits to its exponent, giving it the same dynamic range as full 32-bit precision — the same ability to represent very large and very small numbers — while using only 7 mantissa bits instead of FP32’s 23, sacrificing precision (the number of significant digits) rather than range.

Why This Tradeoff Works for Deep Learning

Neural network training turns out to be considerably more sensitive to a number’s range than to its precision: gradients and activations can vary across many orders of magnitude during training, and losing that range (as happens with FP16’s narrower exponent) causes values to overflow or underflow to zero. Losing precision, by contrast, mostly just adds small amounts of rounding noise to each calculation — noise that stochastic gradient-based optimization is inherently tolerant of, since it’s already working with noisy gradient estimates from mini-batches.

BF16 vs. FP16 vs. FP32

This makes BF16 more numerically forgiving than FP16 for training, without needing techniques like loss scaling that FP16 training often requires to keep values in range. It uses half the memory of FP32 and enables comparable speedups on modern GPU tensor cores, making it the practical default for pretraining and fine-tuning large models today, while FP32 is reserved for a small number of numerically sensitive operations (like certain normalization calculations) that benefit from full precision even within an otherwise BF16 training run.

BF16 Hardware Support

BF16 support isn’t universal across all hardware generations — it requires tensor cores specifically designed to handle the format, present in NVIDIA’s Ampere architecture (A100) and later, as well as Google’s TPUs, where the format actually originated. Older GPU generations (like NVIDIA’s V100) lack native BF16 tensor core support, which is part of why FP16 remained the standard choice on that older hardware despite BF16’s numerical advantages. This hardware dependency is a practical consideration when choosing training infrastructure: teams working with newer GPU generations default to BF16 for training with minimal downside, while those constrained to older hardware may need to use FP16 along with its associated stability techniques like loss scaling.

For inference specifically, BF16 is less commonly the final deployment format compared to training, since further quantization down to INT8 or 4-bit formats usually offers a better latency and memory tradeoff for serving once training-time numerical stability is no longer a concern.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.