Backpropagation
The algorithm that calculates loss gradients backwards through neural layers to update model weights.
Last reviewed: July 25, 2026
Backpropagation is the algorithm that computes how much each weight in a neural network contributed to its prediction error, and it’s the mechanism that makes training deep neural networks computationally feasible. Without an efficient way to calculate these contributions, training networks with millions or billions of parameters — as every modern LLM has — would be intractable.
How It Works
Training a neural network involves a forward pass, where input data flows through the network to produce a prediction, followed by a loss calculation comparing that prediction to the correct answer. Backpropagation then works backward through the network, layer by layer, applying the chain rule from calculus to compute the gradient of the loss with respect to every weight — essentially answering “if I nudge this specific weight slightly, how much would the loss change?” for every single parameter in the network, in one efficient backward sweep, rather than requiring a separate expensive calculation for each parameter individually.
Why It’s Efficient
The key insight is that gradients for earlier layers can be computed by reusing intermediate calculations from later layers, since the chain rule naturally composes this way — this reuse is what makes backpropagation dramatically faster than the naive alternative of numerically estimating each parameter’s gradient independently, which would require a separate forward pass per parameter and would be computationally impossible at the scale of billion-parameter models.
Where It Fits
Backpropagation computes the gradients; an optimizer (like AdamW) then uses those gradients to actually update the weights. The two are often discussed together but are distinct steps: backpropagation answers “which direction should each weight move,” and the optimizer decides “how far to move it,” incorporating considerations like momentum and adaptive learning rates that go beyond the raw gradient itself.
Automatic Differentiation
Modern deep learning frameworks like PyTorch and JAX implement backpropagation through a more general mechanism called automatic differentiation, which builds a computational graph tracking every operation performed during the forward pass, then automatically applies the chain rule backward through that graph to compute gradients — without a developer needing to manually derive or implement the gradient calculation for each new type of layer or operation they write. This is what makes it practical to experiment with novel architectures: as long as an operation is expressed using the framework’s differentiable building blocks, gradients for it are computed automatically, which has been a major factor in the rapid pace of architectural experimentation in deep learning research over the past decade.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.