Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder

Cosine Similarity

A metric that measures the angular distance between two vectors in a high-dimensional space.

Last reviewed: July 25, 2026

Cosine similarity is the most commonly used metric for comparing vector embeddings in semantic search and retrieval systems, measuring the angle between two vectors rather than the distance between them. It produces a score between -1 and 1 (or 0 and 1 for embeddings with only positive values), where 1 means the vectors point in exactly the same direction, 0 means they’re orthogonal (unrelated), and -1 means they point in opposite directions.

Why Angle, Not Distance

Two embedding vectors can point in a very similar direction — representing similar meaning — while having very different magnitudes, for reasons related to how the embedding model was trained rather than the underlying meaning of the text. A distance-based metric like Euclidean distance would penalize this magnitude difference even when the semantic content is nearly identical. Cosine similarity ignores magnitude entirely and considers only direction, making it more robust to this kind of variation, which is why most text embedding models are explicitly trained with cosine similarity in mind as the intended comparison metric.

Calculation

Cosine similarity is calculated as the dot product of two vectors divided by the product of their magnitudes (their Euclidean lengths). When embeddings are pre-normalized to unit length — a common preprocessing step — this calculation simplifies to just the dot product, which is faster to compute at scale and is why many vector databases normalize embeddings on ingestion specifically to enable this optimization.

Practical Relevance

When configuring a vector database index, choosing cosine similarity as the distance metric only produces good results if it matches how the embedding model was actually trained — using a mismatched metric (like raw Euclidean distance on a model trained for cosine similarity) can silently degrade search quality without an obvious error, which is why checking an embedding model’s documentation for its intended similarity metric is an important, easy-to-skip step.

Cosine Similarity Thresholds Are Model-Specific

A common mistake when working with cosine similarity scores is assuming a fixed threshold (like “0.8 means relevant”) transfers across different embedding models — in practice, the distribution of cosine similarity scores an embedding model produces depends heavily on how it was trained, and what counts as a “high” or “meaningfully similar” score for one model can be entirely different for another. Teams building retrieval systems generally need to empirically calibrate similarity thresholds against their specific embedding model and dataset, using labeled examples of genuinely relevant versus irrelevant pairs, rather than assuming a threshold that worked well with one model will transfer directly to another.

Advertisement (In-Content)

Historical figures and technical concepts for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Official Documentation.