Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

FP16 Precision

A half-precision 16-bit floating-point data format providing wide dynamic range for deep learning.

Last reviewed: July 25, 2026

FP16, or half-precision floating point, is a 16-bit numeric format used to store model weights and perform calculations during neural network training and inference, in contrast to the 32-bit FP32 format that was the default in earlier deep learning. Using FP16 instead of FP32 roughly halves the memory required to store a model’s weights and activations, and modern GPUs can execute FP16 matrix multiplications significantly faster than FP32 ones, since GPU tensor cores are specifically optimized for lower-precision arithmetic.

The Tradeoff: Range vs. Precision

FP16 allocates its 16 bits as 1 sign bit, 5 exponent bits, and 10 mantissa bits. The 5 exponent bits give it a much narrower representable range than FP32’s 8 exponent bits — FP16 can only represent numbers roughly between 6×10⁻⁵ and 65,504 before over- or under-flowing. During training, gradients and activations can easily fall outside this range, causing values to silently become zero (underflow) or infinity (overflow), which is why pure FP16 training requires techniques like loss scaling to keep values within the representable range.

FP16 vs. BF16

This range limitation is the main reason BF16 (Brain Floating Point) has become the more common choice for LLM training on modern hardware: BF16 uses the same 8 exponent bits as FP32 (matching its range) while sacrificing more mantissa precision, making it more numerically forgiving during training even though it has fewer significant digits than FP16. FP16 remains widely used for inference, however, where its narrower range is less of a practical problem and its computational speed advantage is fully realized.

Practical Guidance for Choosing a Precision Format

For training, BF16 is now the near-universal default for new model development because of its wider dynamic range, while FP16 remains relevant mainly for inference workloads and for compatibility with older training pipelines and hardware that predates widespread BF16 tensor core support. For inference specifically, the choice between FP16 and further-quantized formats like INT8 or 4-bit typically comes down to how much accuracy loss an application can tolerate in exchange for lower memory usage and higher throughput — FP16 sits at a relatively safe midpoint, offering meaningful memory savings over FP32 with minimal measurable accuracy impact for most models, which is why it remains a common inference default even as more aggressive quantization schemes have become popular for cost-sensitive deployments.

FP16 also remains the format most GPU vendors optimize their oldest generations of tensor cores for, which is part of why inference frameworks targeting older, more widely available hardware often default to FP16 rather than assuming BF16 support is universal.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.