Layer Normalization
A regularization technique that normalizes activations across a single token's feature dimensions.
Last reviewed: July 25, 2026
Layer normalization is a technique used throughout transformer architectures to stabilize training by normalizing the activations within each layer to have a consistent mean and variance, independent of batch size. It’s applied to every transformer block in virtually every modern LLM, making it one of the most structurally important — if easy to overlook — components of the architecture.
How It Differs from Batch Normalization
An earlier normalization technique, batch normalization, normalizes activations across all examples in a training batch for a given feature — which works well for image models but breaks down for sequence models like transformers, where batch statistics can be unstable with variable-length sequences and small batch sizes, and where the technique behaves inconsistently between training and inference. Layer normalization instead normalizes across the feature dimension for a single token independently of other tokens or examples in the batch, making it well suited to variable-length text sequences and consistent between training and inference.
Why It’s Needed
As data passes through many stacked transformer layers, activation values can drift to very large or very small magnitudes, a phenomenon that destabilizes gradient-based training and can cause exploding or vanishing gradients. By renormalizing activations at each layer, layer normalization keeps values in a consistent, well-behaved range throughout the network’s depth, which in practice allows much deeper networks to train reliably.
Where It Sits in the Architecture
Most modern transformers use “pre-norm” placement, applying layer normalization before each attention and feed-forward sublayer rather than after (the original “post-norm” design used in the first Transformer paper). Pre-norm has been found to produce more stable training for very deep networks, which is part of why it became the standard as models scaled to dozens or hundreds of layers. A related variant, RMSNorm, simplifies the calculation by only normalizing variance (not mean) and is used in several recent model families like Llama for a small compute efficiency gain.
RMSNorm as the Modern Default
RMSNorm (Root Mean Square Layer Normalization) has become the more common choice in recent large model architectures, including Llama and several other prominent open-weight families, specifically because it drops the mean-centering step that standard layer normalization performs, only normalizing by the root-mean-square of the activations. This simplification reduces the computation required at each normalization step without measurably hurting model quality in practice, a small but meaningful efficiency gain that compounds across the many normalization operations performed throughout a deep transformer’s forward and backward passes during both training and inference.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.