← Math in Practice
Concept Math behind AI

High-dimensional intuition

When every candidate looks close, inspect the space and benchmark the search contract.

A support-search migration moves a corpus to a new embedding model. The nearest document is still plausible, yet irrelevant notes now crowd the top results. The team asks whether adding dimensions made “nearest” meaningless, or whether preprocessing, corpus, and index behavior changed.

The judgment to keep

In many high-dimensional data sets, distances can become more alike, which weakens raw distance contrast. That observation depends on the data distribution and representation; it does not make every embedding space useless. Inspect labeled neighbors and measure exact and approximate retrieval on the same queries.

TypeScriptGo Distance concentration · cosine · Euclidean · nearest neighbors · ANN evaluation
01 / Read the retrieval failure

A nearest neighbor is relative to a corpus, representation, and metric.

The team compares a query against help-center article vectors. Their old index ranked a small set of relevant articles clearly. After the migration, exact search gives several candidates with similar distances, while approximate search returns a different top five. The phrase “high-dimensional curse” is tempting, but it is not a diagnosis.

Embeddings may have hundreds or thousands of coordinates. An individual coordinate is not usually a stable human-readable feature. Geometry is still useful: vector norms, angles, distances, and ranking margins can reveal a change. But semantic relevance must be checked with labeled queries, and search behavior also depends on corpus membership, filtering, model version, and index settings.

Case file / Search migrationWhy did useful neighbors lose separation?
Observed
Top-k results changed after a model and index migration.
Competing causes
Embedding space, normalization, corpus, filters, metric, or approximate-index recall.
Baseline
Same frozen query set, same corpus snapshot, exact neighbors and judged relevance.
Decision
Do not add dimensions or tune a threshold until the changed stage is isolated.
Evidence to keep: model/version, dimensions, metric, normalization, corpus snapshot, exact top-k, approximate top-k, relevance labels, and query latency.
02 / See what dimension changes

In many random spaces, distances bunch toward a typical scale.

Imagine vectors whose coordinates are drawn independently from a similar, zero-centered distribution. Squared Euclidean distance adds coordinate-wise squared differences. As coordinates accumulate, its total grows roughly with dimension, while relative fluctuations often shrink. Distances can cluster around a typical value, so the nearest and farthest candidates may differ less as a fraction of that scale.

For unit vectors, cosine similarity is the dot product. Under an idealized isotropic random model, unrelated directions tend to have cosine near zero, with spread that typically narrows as dimension grows. Real embeddings are not guaranteed to be isotropic, independent, or random. Training intentionally shapes their neighborhoods; anisotropy, clusters, norms, and task structure can dominate.

EuclideanSum coordinate gaps

More coordinates contribute to squared distance.

Relative contrastMay shrink

Nearest/farthest ratio depends on data distribution.

CosineCompare direction

Useful only with a compatible representation and policy.

Empirical checkMeasure your corpus

Inspect norms, score histograms, and labeled neighbor quality.

03 / Check what “near” means

Cosine and Euclidean answer related but distinct questions.

Euclidean distance includes both orientation and vector length. Cosine similarity compares direction after dividing by magnitudes. If all nonzero vectors are normalized to unit length, then ‖a−b‖² = 2(1−cos(a,b)), so cosine ranking and Euclidean ranking agree. Without that normalization, magnitude can alter Euclidean distance while cosine stays unchanged for positive scaling.

Even with a chosen metric, “nearest” only means nearest under that metric in the indexed representation. It does not mean relevant, correct, safe, or useful. Check whether the vector model and corpus share one coordinate space and whether query and document preprocessing are compatible.

MetricExact definition

Cosine similarity or Euclidean distance, recorded with the index.

NormalizationSame policy

Apply consistently if metric equivalence is expected.

LabelsHuman judgment

Evaluate relevant candidates separately from numeric closeness.

Top-kRanking artifact

Inspect recall, duplicates, and score margins.

04 / Find where intuition breaks

Distance concentration is a lens for experiments, not a verdict on embeddings.

“Curse of dimensionality” describes several related difficulties: volume grows rapidly, finite samples sparsely cover a space, and some distance-based algorithms lose useful contrast or cost. It does not say all high-dimensional learning fails. A learned representation can put task-relevant examples into useful neighborhoods even when raw ambient dimension is large.

For diagnosis, compare distributions rather than one nearest score: norms, pairwise distances or cosines, nearest-neighbor margins, duplicate rates, and relevance by query segment. Compare against simple baselines such as lexical retrieval and against exact vector search. If query behavior differs by language, document length, or category, aggregate metrics can mask the failing slice.

05 / Change the dimension

Watch one synthetic neighborhood, then ask what evidence it does not provide.

Interactive lab / Deterministic synthetic vectorsCompare an anchor with 32 pseudo-random candidates.

Euclidean distance: min 0.37, mean 2.79, max 5.40; nearest/farthest ratio 0.069.

Cosine: min -0.935, mean -0.006, max 0.986.

These fixed-seed, roughly centered synthetic vectors help expose a distribution effect. They are not trained embeddings, and a sample of 32 candidates is far too small to predict production retrieval quality.

06 / Measure approximate search

Approximate nearest neighbors trade exactness for operational cost.

Approximate nearest-neighbor (ANN) indexes reduce query work by searching a subset or compressed representation. They can make large retrieval workloads practical, but the useful question is not simply “is ANN faster?” Compare approximate top-k with exact top-k over the same vectors and queries. Recall@k can be defined as the proportion of exact top-k items also returned by ANN; also measure judged relevance because exact nearest vectors may themselves be semantically poor.

Record latency distributions, memory, index-build time, update cost, and recall as you vary the index's search effort. A setting that improves recall may raise latency or memory. Compare across realistic query types and corpus growth. Index parameters and tradeoffs are implementation-specific; use the chosen index's documentation for its controls and definitions.

Exact baselineGround truth for vector ranking

Same metric, vectors, filters, and corpus snapshot.

ANN recall@kOverlap with exact top-k

Not the same as human relevance.

Service costLatency and memory

Include p95, build, and update behavior.

DecisionChoose a measured point

Make quality and operational limits explicit.

07 / Practice in code

Vector helpers should enforce the comparison contract.

The examples compare cosine and Euclidean distance for equal, nonempty vectors; they reject mismatched dimensions, non-finite coordinates, and zero-vector cosine. They are exact calculations over a pair of vectors, not an ANN index or relevance evaluator.

Go
vectors.go
package main

import (
	"errors"
	"math"
)

func Cosine(a, b []float64) (float64, error) {
	if len(a) == 0 || len(a) != len(b) {
		return 0, errors.New("vectors must have equal nonzero dimensions")
	}
	dot, aa, bb := 0.0, 0.0, 0.0
	for i, x := range a {
		y := b[i]
		if math.IsNaN(x) || math.IsInf(x, 0) || math.IsNaN(y) || math.IsInf(y, 0) {
			return 0, errors.New("coordinates must be finite")
		}
		dot += x * y
		aa += x * x
		bb += y * y
	}
	if aa == 0 || bb == 0 {
		return 0, errors.New("cosine is undefined for a zero vector")
	}
	return dot / math.Sqrt(aa*bb), nil
}

func Euclidean(a, b []float64) (float64, error) {
	if len(a) == 0 || len(a) != len(b) {
		return 0, errors.New("vectors must have equal nonzero dimensions")
	}
	sum := 0.0
	for i, x := range a {
		y := b[i]
		if math.IsNaN(x) || math.IsInf(x, 0) || math.IsNaN(y) || math.IsInf(y, 0) {
			return 0, errors.New("coordinates must be finite")
		}
		d := x - y
		sum += d * d
	}
	return math.Sqrt(sum), nil
}
Transfer: build an exact baseline first, then quantify ANN recall and judged quality against query latency and memory on a representative corpus.
Read the design in your languages.

These choices apply to the comparisons throughout this story.