A traffic spike is a symptom; the overloaded resource is still a question.
At 10:15, the report API’s p95 rises from 240 ms to 1.8 seconds. The database connection pool is near its configured maximum, queue depth is climbing, and one customer’s import produces a visible burst of requests. Those observations make request pressure plausible; they do not yet show that all callers are equally expensive or that the importing customer is the only source of load.
For this teaching scenario, a load test with the same read/write mix sustained about 120 request-cost units per second while meeting the API’s latency target. The test does not prove that every production minute has that exact capacity. Payload size, cache state, query mix, and dependency health can move the safe operating point.
- Observed
- p95 reaches 1.8 s; database queue depth rises.
- Capacity clue
- One comparable test sustained 120 cost units/s.
- Traffic clue
- One tenant has a large import burst.
- Unknown
- Which request types and callers consume the constrained work?
Count useful work at a named boundary before you set a quota.
Use a shared time window and a clear service boundary. Measure attempted and accepted requests, completed work, status codes, queue depth, database pool wait, and latency. Break the data down by route, request class, and authenticated tenant where privacy and cardinality allow. Record retries separately: one user action can become several attempts.
A request count treats a health check and a report that scans millions of rows as equal. If they have meaningfully different cost, use request classes or a documented weight, then validate that the weights track the scarce resource. Do not infer cost from a caller-supplied field that the caller can alter.
Route, tenant, request class, retries.
API admission, worker, pool, or dependency.
Wait time, saturation, errors, or deadline misses.
Compare a scoped cap in a safe rollout.
A bucket separates the average refill rate from the burst it can absorb.
A token bucket holds up to B tokens. It refills at r tokens per second. A request with cost c is admitted when at least c tokens are available, then spends them. The bucket never stores more than B.
If the bucket starts full, its maximum budget over a duration T is B + rT. This is a total allowance, not a promise that the service can perform
that many requests in any arbitrary schedule. Over a long enough period, the sustained
allowance approaches r per second; B is the finite burst credit.
For the scenario, set r = 120 cost units/s and B = 240 cost units. A simultaneous burst of 300 one-unit requests can spend
240 immediately, so 60 are denied (or delayed, if the system queues instead). After one
idle second, 120 tokens refill; the bucket can admit 120 more one-unit requests at that
instant. During a full minute, the theoretical budget is 240 + 120 × 60 = 7,440 units. A steady offer of 200 units/s would present 12,000
units, so at most 7,440 fit this idealized budget and at least 4,560 would be refused or held
back.
Units must agree: if refill is tokens/second, elapsed time is seconds and cost is tokens. A burst of 240 “requests” only means what you think if every request costs one token.
“120 per minute” still leaves a timing rule to choose.
A fixed window counts requests inside aligned buckets such as 10:15:00–10:16:00. It is cheap to reason about, but a caller may send 120 just before the boundary and 120 just after it: 240 requests in roughly two tenths of a second while respecting 120 in each minute bucket.
A sliding window asks how many requests occurred in the trailing interval, for example the previous 60 seconds. It closes that fixed-boundary gap, though exact implementations may store event times or approximate the count with smaller buckets. A token bucket allows short bursts deliberately, then bounds their replenishment. These rules answer different questions; none is automatically the fairest or cheapest for every service.
| Rule | Boundary behavior | Good fit | Watch for |
|---|---|---|---|
| Fixed window | Count resets at aligned clock boundary. | Simple quotas and coarse reporting. | Nearly two windows of work can cluster at an edge. |
| Sliding window | Count events in a moving trailing duration. | A rolling maximum matters. | Exact event storage is expensive; approximations blur edges. |
| Token bucket | Spend burst credit; refill continuously. | Allow bursts while bounding long-run average. | Capacity and refill are separate policy choices. |
- 10:15:59.9
- 120 requests use the first minute’s allowance.
- 10:16:00.1
- 120 requests use the next minute’s allowance.
- Observed interval
- 240 requests in about 0.2 seconds.
- Interpretation
- Correct under the rule; possibly unsafe for the downstream pool.
The key decides whose requests share the same budget.
A global key gives every caller one shared bucket. It protects a single bottleneck, but a busy tenant can consume the budget and block everyone else. A per-user or per-tenant key isolates ordinary callers, but the combined traffic can still exceed a global database capacity. Many systems need both a per-identity fairness limit and a service-wide safety limit.
For authenticated traffic, derive the tenant or account key from trusted server-side identity. A caller-supplied tenant header can be forged. Before authentication, an IP-based key may be a coarse abuse signal, but shared NATs can combine many people and IPv6 address rotation can fragment one actor into many keys. Avoid putting unbounded raw identifiers in metrics labels.
All tenants draw from service capacity.
One import should not consume every tenant’s allowance.
A report export may need a smaller budget than a cached read.
Fairness quotas do not replace a global safety ceiling.
Separate the instant burst from a sustained offer.
This small model starts with an empty bucket before the stated idle period, allows it to refill up to capacity, then presents the burst all at once. The sustained estimate starts with a full bucket and computes a theoretical token budget over the chosen duration. It assumes one token per unit of work and no queue; it does not simulate the exact arrival schedule or prove production capacity.
Starting full makes the sustained budget an upper bound: capacity + refill × duration. Actual admissions depend on when requests arrive and how much each costs. Production policies may queue, reject, or shed work differently.
Refill from elapsed time, then apply the same cost rule.
The examples make three pieces inspectable: token refill, admission by request cost, and
counting timestamps in a trailing interval. The boundary convention in the sliding-window
helper is (now − window, now]; the event exactly at the left edge is
excluded.
Both examples expose refill, request cost, and the trailing-window boundary.
package mathpractice
import (
"errors"
"math"
)
type BucketDecision struct {
Allowed bool
Tokens float64
}
// TakeToken refills a token bucket over elapsed seconds, then attempts one cost.
func TakeToken(tokens, capacity, refillPerSecond, elapsedSeconds, cost float64) (BucketDecision, error) {
values := []float64{tokens, capacity, refillPerSecond, elapsedSeconds, cost}
for _, value := range values {
if math.IsNaN(value) || math.IsInf(value, 0) {
return BucketDecision{}, errors.New("bucket values must be finite")
}
}
if capacity < 0 || refillPerSecond < 0 || elapsedSeconds < 0 || cost <= 0 {
return BucketDecision{}, errors.New("capacity, refill, and elapsed time cannot be negative; cost must be positive")
}
available := math.Min(capacity, math.Max(0, tokens)+refillPerSecond*elapsedSeconds)
if available >= cost {
return BucketDecision{Allowed: true, Tokens: available - cost}, nil
}
return BucketDecision{Allowed: false, Tokens: available}, nil
}
// CountSlidingWindow counts events in (now-windowSeconds, now].
func CountSlidingWindow(timestamps []float64, now, windowSeconds float64) (int, error) {
if math.IsNaN(now) || math.IsInf(now, 0) || math.IsNaN(windowSeconds) || math.IsInf(windowSeconds, 0) || windowSeconds <= 0 {
return 0, errors.New("now must be finite and window must be positive")
}
count := 0
for _, timestamp := range timestamps {
if math.IsNaN(timestamp) || math.IsInf(timestamp, 0) {
return 0, errors.New("timestamps must be finite")
}
if timestamp > now-windowSeconds && timestamp <= now {
count++
}
}
return count, nil
}
// MaximumSpend is the theoretical token budget over a duration from a full bucket.
func MaximumSpend(capacity, refillPerSecond, durationSeconds float64) (float64, error) {
values := []float64{capacity, refillPerSecond, durationSeconds}
for _, value := range values {
if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 {
return 0, errors.New("capacity, refill, and duration must be finite and non-negative")
}
}
return capacity + refillPerSecond*durationSeconds, nil
}
A correct denial still needs a useful contract and evidence.
When a client is over its limit, return a response that identifies the policy outcome and
gives a sensible retry time when the system can calculate one. HTTP 429 Too Many Requests is the standard status for rate limiting; Retry-After can communicate when to try
again. Clients should avoid synchronized immediate retries. Some services queue instead of rejecting,
but the queue still needs a bound and a deadline.
Track allowed and denied work by bounded dimensions such as route and policy, along with queue wait, dependency saturation, latency, and customer-visible errors. Watch for both false acceptance (the resource remains overloaded) and false rejection (healthy callers are blocked). A successful rollout should improve the constrained service while leaving legitimate work within the promised experience.
A single process can keep bucket state in memory. Multiple replicas need shared
coordination, partitioned budgets, or an explicitly approximate scheme. Shared state adds
network latency and can itself fail. If each of n replicas independently
grants a full per-key budget, a caller routed across them may receive close to n times the intended allowance. A partitioned budget can reduce coordination but
needs careful allocation and rebalancing.
Elapsed-time refill depends on a clock. Use monotonic elapsed time within one process so a wall-clock adjustment does not create or erase tokens. Across machines, clocks can skew; do not combine unsynchronized local timestamps as if they were exact. A centralized atomic store can serialize updates but introduces latency and a dependency to monitor. The right choice depends on how strict the quota must be and what happens when coordination is unavailable.
Treat the number as a testable policy, not a permanent constant.
For the report API, the evidence suggests a candidate policy: a weighted per-tenant token bucket with refill near measured sustainable work, a finite burst capacity, plus a global safety ceiling for the database. That is a hypothesis to test, not a final prescription. Start in observation mode or with a narrow rollout. Compare the rejected requests, database wait, queue depth, latency, and customer outcomes with a matched baseline. Adjust the rate and burst separately when the evidence identifies which one is wrong.
If the import is useful but bursty, a queue or asynchronous job may be a better product contract than rejection. If one costly route dominates, a route-specific weight may be better than a uniform request cap. If database contention remains at low API traffic, investigate query plans or locks instead of tightening the limit further.