BM25 Retrieval
A probabilistic term-matching keyword search algorithm widely used in information retrieval.
Last reviewed: July 25, 2026
BM25 (Best Matching 25) is a probabilistic ranking algorithm for keyword-based text search, and despite predating modern vector embeddings by decades, it remains one of the two pillars of production retrieval-augmented generation (RAG) systems alongside dense vector search, most commonly paired together as hybrid search.
How It Works
BM25 scores how relevant a document is to a query based on term frequency — how often query terms appear in the document — balanced against a few refinements that make it more effective than naive term counting. It down-weights terms that appear across many documents in the corpus (since common words are less informative for distinguishing relevance), it dampens the effect of a term appearing many times in one document (so a document isn’t ranked disproportionately higher just for repeating a keyword), and it normalizes for document length, so that longer documents don’t automatically score higher purely by containing more words.
Why It’s Still Used Alongside Vector Search
Vector embeddings excel at matching conceptual meaning, but they can underperform on exact-match queries involving specific identifiers, model numbers, acronyms, or rare technical terms that an embedding model may not represent distinctly from similar-sounding terms. BM25, being a literal term-matching algorithm, has no such weakness — it will always find a document containing an exact keyword match, regardless of whether the embedding model happens to represent that term well.
Where It’s Used
BM25 is the default ranking algorithm in Elasticsearch and OpenSearch, and it’s the “sparse” half of most hybrid search implementations, run in parallel with dense vector retrieval and merged into a combined ranking — typically via Reciprocal Rank Fusion — to get the benefits of both exact and semantic matching in a single RAG pipeline.
BM25’s Tunable Parameters
BM25’s scoring formula includes two tunable parameters that affect how it weighs term frequency and document length: k1 controls how quickly the scoring saturates as a term’s frequency in a document increases (preventing a document from scoring disproportionately higher just for repeating a term many times), and b controls how strongly document length normalization is applied. While most implementations use sensible defaults, tuning these parameters against a representative set of queries and relevance judgments for a specific corpus can meaningfully improve BM25’s ranking quality — a detail often overlooked by teams that treat BM25 as a fixed, non-configurable baseline rather than a scoring function with its own optimization surface.
Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.