Product Quantization (PQ)
A lossy vector compression technique that reduces database RAM footprints.
Last reviewed: July 25, 2026
Product Quantization (PQ) is a vector compression technique used in large-scale vector search systems to dramatically reduce the memory footprint of stored embeddings, at the cost of some retrieval accuracy — it’s the compression component behind the widely used IVF-PQ indexing method.
How It Works
Instead of storing a vector’s full sequence of floating-point values, product quantization splits each vector into several smaller sub-vectors, and for each sub-vector position, learns a small “codebook” of representative values (typically 256 of them, via a clustering algorithm like k-means) from the actual data. Each sub-vector in the dataset is then replaced by the index of its nearest codebook entry — a single byte, since 256 fits in 8 bits — rather than storing the original floating-point sub-vector directly.
Why This Saves So Much Memory
A typical 768-dimensional embedding stored as 32-bit floats requires about 3KB per vector. Splitting that vector into, say, 96 sub-vectors of 8 dimensions each, and replacing each sub-vector with a single codebook index byte, reduces the stored representation to just 96 bytes — roughly a 30x compression ratio. This compounds enormously at scale: a billion-vector index that would require terabytes of memory in raw float form can potentially fit in a fraction of that with product quantization applied.
The Accuracy Tradeoff
Because each sub-vector is approximated by its nearest codebook entry rather than stored exactly, distance calculations between quantized vectors are themselves approximate, meaning searches over a PQ-compressed index have somewhat lower recall (a higher chance of missing the true nearest neighbor) than searching uncompressed vectors. This tradeoff is generally accepted for use cases where the alternative — needing enough memory to store billions of raw vectors — simply isn’t practical or cost-effective.
Combining Product Quantization With Reranking
Because product quantization introduces approximation error into distance calculations, production systems often combine it with a reranking step: an initial PQ-based search retrieves a larger candidate set quickly using compressed vectors, and then a final pass re-scores just those candidates using their original, uncompressed representations (stored separately, often on cheaper storage than the compressed index needs to be held in memory) for a more accurate final ranking. This two-stage pattern — fast approximate search followed by precise reranking on a small candidate set — is a recurring architecture across large-scale vector search systems, appearing in similar form whether the underlying compression technique is product quantization, scalar quantization, or another approximation method entirely.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.