← Math in Practice
Concept Math in engineering decisions

Rate limiting math

Choose a limit from measured capacity and traffic shape, not from a round number.

A report API starts timing out shortly after a customer imports a large batch. Its request rate spiked, but the same endpoint also serves ordinary dashboard reads. The team needs to protect the database without turning a legitimate import into a blanket outage.

The judgment to keep

First locate the constrained work and identify who is producing it. Then choose a limit whose unit, burst allowance, time model, and key match the capacity and fairness question you mean to answer.

TypeScriptGo Token bucket · windows · burst and sustained rate · keys · telemetry
01 / Read the overload

A traffic spike is a symptom; the overloaded resource is still a question.

At 10:15, the report API’s p95 rises from 240 ms to 1.8 seconds. The database connection pool is near its configured maximum, queue depth is climbing, and one customer’s import produces a visible burst of requests. Those observations make request pressure plausible; they do not yet show that all callers are equally expensive or that the importing customer is the only source of load.

For this teaching scenario, a load test with the same read/write mix sustained about 120 request-cost units per second while meeting the API’s latency target. The test does not prove that every production minute has that exact capacity. Payload size, cache state, query mix, and dependency health can move the safe operating point.

Case file / Report APILatency climbs during a customer import.
Observed
p95 reaches 1.8 s; database queue depth rises.
Capacity clue
One comparable test sustained 120 cost units/s.
Traffic clue
One tenant has a large import burst.
Unknown
Which request types and callers consume the constrained work?
Leave with: one suspected bottleneck, at least one competing cause, and the measurements that could distinguish them.
02 / Measure before limiting

Count useful work at a named boundary before you set a quota.

Use a shared time window and a clear service boundary. Measure attempted and accepted requests, completed work, status codes, queue depth, database pool wait, and latency. Break the data down by route, request class, and authenticated tenant where privacy and cardinality allow. Record retries separately: one user action can become several attempts.

A request count treats a health check and a report that scans millions of rows as equal. If they have meaningfully different cost, use request classes or a documented weight, then validate that the weights track the scarce resource. Do not infer cost from a caller-supplied field that the caller can alter.

01 / DemandWho and what?

Route, tenant, request class, retries.

02 / BoundaryWhere does work queue?

API admission, worker, pool, or dependency.

03 / EffectWhat degrades first?

Wait time, saturation, errors, or deadline misses.

04 / TestWhat changes under control?

Compare a scoped cap in a safe rollout.

03 / Model a token bucket

A bucket separates the average refill rate from the burst it can absorb.

A token bucket holds up to B tokens. It refills at r tokens per second. A request with cost c is admitted when at least c tokens are available, then spends them. The bucket never stores more than B.

If the bucket starts full, its maximum budget over a duration T is B + rT. This is a total allowance, not a promise that the service can perform that many requests in any arbitrary schedule. Over a long enough period, the sustained allowance approaches r per second; B is the finite burst credit.

For the scenario, set r = 120 cost units/s and B = 240 cost units. A simultaneous burst of 300 one-unit requests can spend 240 immediately, so 60 are denied (or delayed, if the system queues instead). After one idle second, 120 tokens refill; the bucket can admit 120 more one-unit requests at that instant. During a full minute, the theoretical budget is 240 + 120 × 60 = 7,440 units. A steady offer of 200 units/s would present 12,000 units, so at most 7,440 fit this idealized budget and at least 4,560 would be refused or held back.

Refill after elapsed time Δttokens′ = min(B, tokens + r × Δt)
Admission ruleallow if tokens′ ≥ request cost c
Full-bucket budget over Tmaximum spend = B + rT

Units must agree: if refill is tokens/second, elapsed time is seconds and cost is tokens. A burst of 240 “requests” only means what you think if every request costs one token.

04 / Compare window rules

“120 per minute” still leaves a timing rule to choose.

A fixed window counts requests inside aligned buckets such as 10:15:00–10:16:00. It is cheap to reason about, but a caller may send 120 just before the boundary and 120 just after it: 240 requests in roughly two tenths of a second while respecting 120 in each minute bucket.

A sliding window asks how many requests occurred in the trailing interval, for example the previous 60 seconds. It closes that fixed-boundary gap, though exact implementations may store event times or approximate the count with smaller buckets. A token bucket allows short bursts deliberately, then bounds their replenishment. These rules answer different questions; none is automatically the fairest or cheapest for every service.

Three common meanings of “120 per minute”
RuleBoundary behaviorGood fitWatch for
Fixed windowCount resets at aligned clock boundary.Simple quotas and coarse reporting.Nearly two windows of work can cluster at an edge.
Sliding windowCount events in a moving trailing duration.A rolling maximum matters.Exact event storage is expensive; approximations blur edges.
Token bucketSpend burst credit; refill continuously.Allow bursts while bounding long-run average.Capacity and refill are separate policy choices.
Boundary trace / 120 per minuteTwo legal fixed windows, one sharp burst.
10:15:59.9
120 requests use the first minute’s allowance.
10:16:00.1
120 requests use the next minute’s allowance.
Observed interval
240 requests in about 0.2 seconds.
Interpretation
Correct under the rule; possibly unsafe for the downstream pool.
05 / Choose the limit key

The key decides whose requests share the same budget.

A global key gives every caller one shared bucket. It protects a single bottleneck, but a busy tenant can consume the budget and block everyone else. A per-user or per-tenant key isolates ordinary callers, but the combined traffic can still exceed a global database capacity. Many systems need both a per-identity fairness limit and a service-wide safety limit.

For authenticated traffic, derive the tenant or account key from trusted server-side identity. A caller-supplied tenant header can be forged. Before authentication, an IP-based key may be a coarse abuse signal, but shared NATs can combine many people and IPv6 address rotation can fragment one actor into many keys. Avoid putting unbounded raw identifiers in metrics labels.

01 / GlobalProtect a shared pool

All tenants draw from service capacity.

02 / TenantKeep one tenant fair

One import should not consume every tenant’s allowance.

03 / RouteMatch different costs

A report export may need a smaller budget than a cached read.

04 / CombinedCheck both constraints

Fairness quotas do not replace a global safety ceiling.

06 / Change the assumptions

Separate the instant burst from a sustained offer.

This small model starts with an empty bucket before the stated idle period, allows it to refill up to capacity, then presents the burst all at once. The sustained estimate starts with a full bucket and computes a theoretical token budget over the chosen duration. It assumes one token per unit of work and no queue; it does not simulate the exact arrival schedule or prove production capacity.

Tokens before burst120
Admitted in burst120
Denied or delayed180
Offered units12,000
Maximum token budget7,440
Estimated not admitted4,560

Starting full makes the sustained budget an upper bound: capacity + refill × duration. Actual admissions depend on when requests arrive and how much each costs. Production policies may queue, reject, or shed work differently.

07 / Practice in code

Refill from elapsed time, then apply the same cost rule.

The examples make three pieces inspectable: token refill, admission by request cost, and counting timestamps in a trailing interval. The boundary convention in the sliding-window helper is (now − window, now]; the event exactly at the left edge is excluded.

Compare the same rate-limit calculations in TypeScript and Go.

Both examples expose refill, request cost, and the trailing-window boundary.

GoToken bucket and sliding window helpers
rate-limits.go
package mathpractice

import (
	"errors"
	"math"
)

type BucketDecision struct {
	Allowed bool
	Tokens  float64
}

// TakeToken refills a token bucket over elapsed seconds, then attempts one cost.
func TakeToken(tokens, capacity, refillPerSecond, elapsedSeconds, cost float64) (BucketDecision, error) {
	values := []float64{tokens, capacity, refillPerSecond, elapsedSeconds, cost}
	for _, value := range values {
		if math.IsNaN(value) || math.IsInf(value, 0) {
			return BucketDecision{}, errors.New("bucket values must be finite")
		}
	}
	if capacity < 0 || refillPerSecond < 0 || elapsedSeconds < 0 || cost <= 0 {
		return BucketDecision{}, errors.New("capacity, refill, and elapsed time cannot be negative; cost must be positive")
	}
	available := math.Min(capacity, math.Max(0, tokens)+refillPerSecond*elapsedSeconds)
	if available >= cost {
		return BucketDecision{Allowed: true, Tokens: available - cost}, nil
	}
	return BucketDecision{Allowed: false, Tokens: available}, nil
}

// CountSlidingWindow counts events in (now-windowSeconds, now].
func CountSlidingWindow(timestamps []float64, now, windowSeconds float64) (int, error) {
	if math.IsNaN(now) || math.IsInf(now, 0) || math.IsNaN(windowSeconds) || math.IsInf(windowSeconds, 0) || windowSeconds <= 0 {
		return 0, errors.New("now must be finite and window must be positive")
	}
	count := 0
	for _, timestamp := range timestamps {
		if math.IsNaN(timestamp) || math.IsInf(timestamp, 0) {
			return 0, errors.New("timestamps must be finite")
		}
		if timestamp > now-windowSeconds && timestamp <= now {
			count++
		}
	}
	return count, nil
}

// MaximumSpend is the theoretical token budget over a duration from a full bucket.
func MaximumSpend(capacity, refillPerSecond, durationSeconds float64) (float64, error) {
	values := []float64{capacity, refillPerSecond, durationSeconds}
	for _, value := range values {
		if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 {
			return 0, errors.New("capacity, refill, and duration must be finite and non-negative")
		}
	}
	return capacity + refillPerSecond*durationSeconds, nil
}
Boundary check: with capacity 240, refill 120 tokens/s, 1 second elapsed, and cost 1, the bucket refills to 240 (it cannot exceed capacity) and admits one request, leaving 239 tokens.
08 / Read the response

A correct denial still needs a useful contract and evidence.

When a client is over its limit, return a response that identifies the policy outcome and gives a sensible retry time when the system can calculate one. HTTP 429 Too Many Requests is the standard status for rate limiting; Retry-After can communicate when to try again. Clients should avoid synchronized immediate retries. Some services queue instead of rejecting, but the queue still needs a bound and a deadline.

Track allowed and denied work by bounded dimensions such as route and policy, along with queue wait, dependency saturation, latency, and customer-visible errors. Watch for both false acceptance (the resource remains overloaded) and false rejection (healthy callers are blocked). A successful rollout should improve the constrained service while leaving legitimate work within the promised experience.

A single process can keep bucket state in memory. Multiple replicas need shared coordination, partitioned budgets, or an explicitly approximate scheme. Shared state adds network latency and can itself fail. If each of n replicas independently grants a full per-key budget, a caller routed across them may receive close to n times the intended allowance. A partitioned budget can reduce coordination but needs careful allocation and rebalancing.

Elapsed-time refill depends on a clock. Use monotonic elapsed time within one process so a wall-clock adjustment does not create or erase tokens. Across machines, clocks can skew; do not combine unsynchronized local timestamps as if they were exact. A centralized atomic store can serialize updates but introduces latency and a dependency to monitor. The right choice depends on how strict the quota must be and what happens when coordination is unavailable.

09 / Make the next call

Treat the number as a testable policy, not a permanent constant.

For the report API, the evidence suggests a candidate policy: a weighted per-tenant token bucket with refill near measured sustainable work, a finite burst capacity, plus a global safety ceiling for the database. That is a hypothesis to test, not a final prescription. Start in observation mode or with a narrow rollout. Compare the rejected requests, database wait, queue depth, latency, and customer outcomes with a matched baseline. Adjust the rate and burst separately when the evidence identifies which one is wrong.

If the import is useful but bursty, a queue or asynchronous job may be a better product contract than rejection. If one costly route dominates, a route-specific weight may be better than a uniform request cap. If database contention remains at low API traffic, investigate query plans or locks instead of tightening the limit further.

Take the idea with you: name the constrained resource, the measured unit, the refill rate, the burst allowance, the identity key, the denial behavior, and the telemetry that would make you revise the policy.