Activation Function
A mathematical formula introducing non-linear thresholds to model outputs, allowing networks to learn complex representations.
Last reviewed: July 25, 2026
An activation function is the non-linear transformation applied to a neuron’s output within a neural network, and it’s the single ingredient that allows deep learning to model complex, non-linear relationships rather than just stacking linear operations. Without activation functions, no matter how many layers a network had, the entire stack would mathematically collapse into a single equivalent linear transformation — layers of matrix multiplication compose into another matrix multiplication, gaining no additional representational power from depth.
Common Activation Functions
Early neural networks commonly used sigmoid or tanh functions, which squash outputs into a bounded range but suffer from vanishing gradients — for inputs far from zero, their gradient approaches zero, which slows or stalls learning in deep networks. ReLU (Rectified Linear Unit), which simply outputs zero for negative inputs and passes positive inputs through unchanged, became the standard for most deep networks because it avoids this vanishing gradient problem for positive inputs and is cheap to compute.
Modern transformer architectures typically use smoother variants like GELU (Gaussian Error Linear Unit) or SwiGLU, which behave similarly to ReLU but transition more smoothly around zero, and have been empirically found to produce better results in large language model training, at a modest additional compute cost.
Why It Matters
The choice of activation function is a real architectural decision with measurable effects on training stability and final model quality — it’s one of the details that differs between model families (Llama uses SwiGLU, for example) and is generally not something practitioners change when fine-tuning, since it’s baked into the pretrained model’s architecture.
How Activation Choice Affects Practical Fine-Tuning
While practitioners fine-tuning an existing pretrained model generally don’t change its activation function — it’s baked into the architecture and changing it would require retraining from scratch — understanding which activation a model uses can matter when implementing custom layers or adapters (like LoRA) on top of it, since compatibility with the surrounding network’s numerical behavior depends on matching conventions. It’s also relevant when comparing benchmark results across model families: differences in reported performance between otherwise similarly sized models can sometimes be partially attributed to architectural choices like activation function alongside more commonly cited factors like training data and parameter count.
One notable exception worth knowing about is when a team distills or prunes a model into a new architecture, at which point activation function choice becomes a live design decision again rather than an inherited constraint from the base model.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.