A ranking score is the beginning of an investigation, not its conclusion.
A help-center search service turns a query and each article into vectors, then compares the query vector with candidate vectors. The highest-ranked article sounds relevant, but the engineer's task is to understand the score well enough to check the behavior. We will use tiny two-coordinate vectors so every operation can be verified by hand. Real embedding vectors often have hundreds or thousands of coordinates, and those coordinates usually do not have simple human-readable labels such as “password” or “billing.”
- Query
- “reset a forgotten password” encoded as q = (2, 1).
- Candidate
- A password-reset article encoded as d = (3, 4).
- Observed score
- q · d = 2×3 + 1×4 = 10.
- Question
- What changes if a vector's direction or length changes?
A vector is an ordered list; positions have meaning only within the chosen representation.
The query vector q = (2, 1) has two coordinates: 2 in position one and 1 in
position two. The candidate d = (3, 4) also has two coordinates. Position one must
refer to the same learned dimension in both vectors for a coordinate-by-coordinate comparison
to make sense. If one vector comes from a different model version or coordinate ordering, matching
positions may no longer be comparable.
In a hand-built example, we can imagine axes such as “password reset intent” and “account access.” That picture helps introduce the arithmetic, but deployed embeddings are learned representations. An individual dimension generally does not carry a stable, plain-English label. Meaning is represented by patterns across coordinates, relative to the model that produced them.
Two ordered coordinates.
Same dimension and coordinate order.
Multiply corresponding positions.
One scalar score from two vectors.
Pair matching coordinates, multiply, and sum.
For vectors a = (a₁, a₂, …, aₙ) and b = (b₁, b₂, …, bₙ), the dot
product is a · b = Σ aᵢbᵢ. In two dimensions, the full calculation is:
(2, 1) · (3, 4) = (2 × 3) + (1 × 4) = 6 + 4 = 10
The first product contributes 6 and the second contributes 4. Adding them gives the scalar 10. The vectors themselves are still two-dimensional; the dot product is a single number computed from them. This is also the pattern used with longer vectors: align each coordinate pair, multiply, and add the products.
| Position | Query q | Candidate d | Product |
|---|---|---|---|
| 1 | 2 | 3 | 2 × 3 = 6 |
| 2 | 1 | 4 | 1 × 4 = 4 |
| Sum | Dot product q · d | 6 + 4 = 10 | |
The calculation also has an algebraic role: it measures how much one vector projects along the other, scaled by their lengths. That relationship is why direction and magnitude both matter. We will separate them next rather than reading the raw score as a direction-only comparison.
The same direction can produce different dot products when vector lengths differ.
The geometric identity a · b = |a| |b| cos(θ) connects the dot product to
vector magnitudes and the angle between them. When vectors point in a similar direction,
cos(θ) is positive; when they are perpendicular it is zero; when they point in opposite
directions it is negative. But the lengths |a| and |b| scale the result.
For q = (2, 1) and d = (3, 4), the magnitudes are |q| = √(2² + 1²) = √5 ≈ 2.236 and |d| = √(3² + 4²) = 5. The cosine of the angle is therefore 10 ÷ (√5 × 5) ≈ 0.894. A raw dot product of 10 includes those lengths; cosine
similarity divides them out and compares orientation for nonzero vectors.
If we double the candidate to (6, 8), its direction is unchanged but the dot
product becomes (2×6) + (1×8) = 20. Its cosine similarity with the query
remains about 0.894. This is a useful check when you need to know whether a score is
sensitive to vector length or intended to focus on direction.
Raw score for these coordinates.
Both vectors have nonzero magnitude.
Dot divided by the product of magnitudes.
“10” is a score in a particular system, not a universal rating scale.
Do not carry a threshold across models without checking it.
A new encoder, changed normalization, or new corpus can shift the score distribution. The ordering may change, score magnitudes may change, or both. Record which model version and metric produced a score, then compare a labeled query set before deciding that the old cutoff still means the same thing.
Watch each coordinate contribute to the total and compare direction separately.
Enter two-coordinate vectors. The trace calculates the dot product, each vector's Euclidean magnitude, and cosine similarity. The latter is shown only when both vectors have nonzero length. These are arithmetic demonstrations, not real embeddings or a relevance model.
Assumption: coordinates are in the same ordered, compatible vector space.
Try d = (6, 8). The dot product doubles from 10 to 20 because the candidate
magnitude doubles; cosine stays about 0.894 because its direction did not change. Try d = (−2, −1): the dot product is −5 and cosine is −1, showing an opposite
direction in this simple coordinate system. Actual model metrics and score ranges still
depend on their implementation and data.
Make dimension checks and the scoring rule visible.
The examples calculate a dot product, Euclidean magnitude, and cosine for the worked vectors. They reject empty or different-sized vectors and non-finite coordinate values. The snippets use small arrays to make the calculation readable; production embedding pipelines also need consistent model versions, data handling, and evaluation.
Both examples use q = (2, 1) and d = (3, 4): dot = 10 and cosine ≈ 0.894.
export type Vector = readonly number[];
/** Dot product for finite, equal-length vectors. */
export function dot(left: Vector, right: Vector): number {
if (left.length === 0 || left.length !== right.length) {
throw new RangeError('Vectors must have the same nonzero dimension.');
}
let total = 0;
for (let i = 0; i < left.length; i += 1) {
const a = left[i];
const b = right[i];
if (!Number.isFinite(a) || !Number.isFinite(b)) {
throw new RangeError('Vector coordinates must be finite.');
}
total += a * b;
}
if (!Number.isFinite(total)) throw new RangeError('Dot product overflowed the number range.');
return total;
}
/** Euclidean length (L2 norm) of a finite, nonempty vector. */
export function magnitude(vector: Vector): number {
if (vector.length === 0 || !vector.every(Number.isFinite)) {
throw new RangeError('Vector must have finite coordinates and nonzero dimension.');
}
let result = 0;
for (const value of vector) result = Math.hypot(result, value);
if (!Number.isFinite(result)) throw new RangeError('Magnitude is outside the number range.');
return result;
}
/** Cosine similarity for nonzero vectors in the same coordinate space. */
export function cosineSimilarity(left: Vector, right: Vector): number {
const denominator = magnitude(left) * magnitude(right);
if (denominator === 0 || !Number.isFinite(denominator)) {
throw new RangeError('Cosine similarity requires finite, nonzero magnitudes.');
}
return dot(left, right) / denominator;
}
const query = [2, 1] as const;
const candidate = [3, 4] as const;
console.log(dot(query, candidate)); // 10
console.log(magnitude(query).toFixed(3)); // 2.236
console.log(cosineSimilarity(query, candidate).toFixed(3)); // 0.894
package main
import (
"errors"
"fmt"
"math"
)
// Dot returns the dot product of finite, equal-length vectors.
func Dot(left, right []float64) (float64, error) {
if len(left) == 0 || len(left) != len(right) {
return 0, errors.New("vectors must have the same nonzero dimension")
}
var total float64
for i := range left {
if math.IsNaN(left[i]) || math.IsInf(left[i], 0) || math.IsNaN(right[i]) || math.IsInf(right[i], 0) {
return 0, errors.New("vector coordinates must be finite")
}
total += left[i] * right[i]
}
if math.IsNaN(total) || math.IsInf(total, 0) {
return 0, errors.New("dot product overflowed the number range")
}
return total, nil
}
// Magnitude returns the Euclidean length of a finite, nonempty vector.
func Magnitude(vector []float64) (float64, error) {
if len(vector) == 0 {
return 0, errors.New("vector must have nonzero dimension")
}
result := 0.0
for _, value := range vector {
if math.IsNaN(value) || math.IsInf(value, 0) {
return 0, errors.New("vector coordinates must be finite")
}
result = math.Hypot(result, value)
}
if math.IsInf(result, 0) || math.IsNaN(result) {
return 0, errors.New("magnitude is outside the number range")
}
return result, nil
}
// CosineSimilarity compares direction and requires nonzero vector magnitudes.
func CosineSimilarity(left, right []float64) (float64, error) {
dot, err := Dot(left, right)
if err != nil {
return 0, err
}
leftMagnitude, err := Magnitude(left)
if err != nil {
return 0, err
}
rightMagnitude, err := Magnitude(right)
if err != nil {
return 0, err
}
denominator := leftMagnitude * rightMagnitude
if denominator == 0 || math.IsInf(denominator, 0) {
return 0, errors.New("cosine similarity requires finite, nonzero magnitudes")
}
return dot / denominator, nil
}
func main() {
query := []float64{2, 1}
candidate := []float64{3, 4}
dot, err := Dot(query, candidate)
if err != nil {
panic(err)
}
queryMagnitude, _ := Magnitude(query)
candidateMagnitude, _ := Magnitude(candidate)
cosine, err := CosineSimilarity(query, candidate)
if err != nil {
panic(err)
}
fmt.Printf("dot=%.0f query-length=%.3f candidate-length=%.3f cosine=%.3f\n", dot, queryMagnitude, candidateMagnitude, cosine)
}
- Google Machine Learning Crash Course: Similarity measures — dot product and cosine similarity as related but distinct comparison measures.
- NumPy: numpy.dot — dot products for one-dimensional arrays and related higher-dimensional operations.
Before changing a threshold, determine which part of the retrieval path changed.
- Reproduce. Save the exact query, returned candidates, and index/model versions.
- Inspect representation. Confirm dimensions, coordinate order, normalization, missing values, and stale vectors.
- Recompute a sample. Trace products for a tiny vector by hand or compare a known test vector against the service result.
- Compare decisions. Evaluate rankings and any threshold on representative, labeled queries rather than one example.
- Check transfer. If you switch to cosine or a new encoder, recalibrate from measured relevance outcomes.
The product begins normalizing every vector before ranking.
For nonzero vectors, what happens to the dot product after each vector is normalized to length 1? Explain why the resulting score now matches cosine similarity, then name two checks you would run before carrying an old relevance threshold into this changed pipeline.