Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

Multi-Query Attention (MQA)

An attention mechanism where all query heads share a single key-value head to minimize KV cache storage overhead.

Last reviewed: July 25, 2026

Technical Overview of MQA

Multi-Query Attention (MQA) is an attention mechanism where all Query heads share a single Key-Value (KV) projection head.

During LLM generation, storing the Key-Value projection vectors for preceding tokens (the KV cache) consumes significant GPU memory, limiting concurrent serving capacity. MQA addresses this memory bottleneck by forcing all query heads to share a single Key-Value head, reducing the memory footprint of the KV cache at the cost of a slight reduction in model reasoning capacity.

Key Architecture & Implementation

In Multi-Head Attention (MHA), each query head has a dedicated Key and Value head. In contrast, MQA projects the Key and Value matrices only once:

  • Query Projection Shape: [Batch,SeqLen,HeadsQ,Dim][Batch, SeqLen, Heads_Q, Dim]
  • Key/Value Projection Shape: [Batch,SeqLen,1,Dim][Batch, SeqLen, 1, Dim]

During the attention step, the single KV representation is broadcast across all Query heads to calculate attention logits.

[ Query Head 1 ] ---\
[ Query Head 2 ] ----+---> [ Shared KV Head ]
[ Query Head 3 ] ---/

Core Parameters

  • Memory Savings: Reduces the VRAM footprint of the KV cache by a factor equal to the number of query heads (e.g., a 32x reduction for a model with 32 query heads).
  • Throughput Increase: Frees up VRAM, allowing larger batch sizes and higher throughput during serving.

Real-world Applications

  • Used in early LLMs like Falcon and Google’s PaLM to optimize serving efficiency.

The Tradeoff Against Standard Multi-Head Attention

In standard multi-head attention, every attention head has its own separate key and value projections, meaning the KV cache grows proportionally to the number of heads. Multi-Query Attention has all heads share a single key-value pair, shrinking the KV cache by roughly the number of heads used — for a model with 32 attention heads, that’s up to a 32x reduction in KV cache memory per token, which directly translates to serving more concurrent users or longer contexts on the same GPU. The cost is some loss of representational capacity, since all query heads are now constrained to attend using the same shared keys and values rather than each learning its own distinct projection, which can measurably hurt model quality compared to full multi-head attention at the same parameter count. This tradeoff is why Grouped-Query Attention (GQA) has become the more common middle-ground choice in recent model families like Llama 2 and 3 — it shares key-value heads across small groups of query heads rather than sharing a single pair across all of them, recovering much of MQA’s memory savings while preserving more of standard attention’s quality.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.