Inverted File Index (IVF)
A vector search optimization that clusters vector spaces to limit search scopes.
Last reviewed: July 25, 2026
An Inverted File index (IVF) is a vector search structure that partitions a large collection of embeddings into clusters, so that a query only needs to be compared against vectors in a small number of relevant clusters rather than the entire dataset — the same fundamental idea behind classic keyword search inverted indexes, adapted for vector similarity.
How It Works
During index construction, IVF runs a clustering algorithm (typically k-means) over the full set of vectors to identify a fixed number of cluster centroids, then assigns each vector in the dataset to its nearest centroid. At query time, the query vector is first compared against the (much smaller) set of centroids to identify the closest few clusters — a parameter usually called nprobe controls how many clusters get searched — and only vectors within those selected clusters are compared directly against the query.
Why It’s Efficient
If a dataset has 1,000 clusters and a search only probes the 10 closest ones, the search only needs to examine roughly 1% of the total vectors directly, a substantial reduction compared to a brute-force scan of every vector. The tradeoff is that if the true nearest neighbor happens to sit in a cluster that wasn’t among the ones probed — which can happen for vectors sitting near a cluster boundary — the search will simply miss it, meaning IVF search is approximate and its accuracy (recall) can be tuned by adjusting how many clusters are probed at query time.
IVF-PQ
IVF is frequently combined with product quantization (as IVF-PQ) to additionally compress the vectors stored within each cluster, since the two techniques address different aspects of the scaling problem: IVF reduces how many vectors need to be compared per query, while PQ reduces the memory footprint of each vector being compared.
Choosing the Number of Clusters
The number of clusters used when building an IVF index is itself a meaningful tuning parameter: too few clusters means each cluster contains a large number of vectors, reducing the search speedup IVF provides; too many clusters means the initial coarse search step (comparing the query against cluster centroids) itself becomes slower, and individual clusters may become so small that the clustering’s structure stops meaningfully narrowing the search space. A common rule of thumb suggests choosing a number of clusters roughly proportional to the square root of the total number of vectors in the dataset, though the truly optimal value depends on the specific data distribution and is often tuned empirically for a given corpus.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.