Feed-Forward Network (FFN)
A fully-connected neural network layer processing token vectors independently after attention calculations.
Last reviewed: July 25, 2026
The feed-forward network (FFN), also called the MLP (multi-layer perceptron) block, is one of the two core sublayers that make up every transformer block, alongside the self-attention mechanism. While attention lets tokens exchange information with each other, the feed-forward network processes each token’s representation independently, applying the same learned transformation to every position in the sequence.
Structure
A transformer’s FFN typically consists of two linear transformations with a non-linear activation function (commonly GELU or SwiGLU) in between: the first linear layer expands the token’s representation to a higher-dimensional space — often 4x the model’s hidden dimension — and the second projects it back down to the original dimension. This expand-then-contract structure gives the network extra representational capacity to process each token’s information before passing it to the next transformer block.
Why It Matters
Research analyzing where a transformer’s parameters and computation actually go has found that feed-forward layers account for the majority of a typical transformer’s total parameter count — often roughly two-thirds — making them the primary place where a model’s factual knowledge and learned patterns are believed to be stored, as opposed to attention layers, which are more associated with routing and relating information between tokens. This is part of why techniques like Mixture of Experts specifically replace the feed-forward layer with multiple specialized expert FFNs and a routing mechanism, since it’s the component with the most capacity to scale.
Relevance to Model Efficiency
Because feed-forward layers dominate parameter count, they’re also a primary target for compression techniques like quantization and pruning, and their computational cost (which scales with sequence length linearly, unlike attention’s quadratic scaling) is a meaningful part of total inference cost for long sequences.
SwiGLU and Modern FFN Variants
Many recent model families (Llama, Mistral, and others) use a specific FFN variant called SwiGLU, which replaces the simple two-layer expand-then-contract structure with a gated design: the input passes through two separate linear projections, one of which is transformed by a SiLU (Swish) activation function and then multiplied element-wise against the other, before a final linear layer projects the result back down to the model’s hidden dimension. This gating mechanism has been empirically found to improve model quality compared to a standard ReLU or GELU-based FFN at the same parameter count, at the cost of requiring three weight matrices instead of two — a tradeoff most modern model architectures have concluded is worthwhile given the resulting quality improvement.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.