The cluster looks comfortable while one service falls behind.
For this illustrative incident, eight API replicas receive about 820 accepted requests each second and complete about 760 requests each second over a five-minute window. The queue grows while p95 latency reaches 1.2 seconds. A prior load test sustained 100 requests/s per replica for this request mix while keeping p95 at or below 250 ms and errors below 0.1%.
The numbers are a teaching scenario, not production measurements. “Accepted” means work enters this service’s queue; completions leave the same boundary. Count retry attempts in offered demand, and track them separately so retries do not masquerade as new customer work.
Demand and capacity need the same unit and the same service boundary.
Demand is work offered to or accepted by the boundary you are studying, here requests per second entering the API queue. Be explicit about which one: attempted ingress can include requests rejected before queueing, while accepted arrivals are the right quantity for the queue balance. Retries are work too, even when they do not represent a new user action.
Capacity is the sustainable rate this configured service can handle while meeting its stated objective for a defined workload. The illustrative test says one replica sustained 100 requests/s for the current request mix with p95 at or below 250 ms and errors below 0.1%. That is a measured operating point, not a universal maximum. Change the payloads, cache hit rate, database, concurrency, or latency target and the estimate may change.
Utilization here is a planning ratio: offered demand divided by measured capacity. A resource gauge such as CPU utilization is a different ratio. Planning utilization can exceed 100% when offered work exceeds the measured rate; that means a deficit in this model, not a CPU gauge that somehow reads 102.5%.
Headroom is capacity minus demand. Positive headroom is room under the chosen measurement; zero means demand equals it; negative headroom is a deficit. Keep the units, requests/s, so the sign and magnitude remain interpretable.
The offered rate is slightly above the tested service rate.
| Quantity | Calculation | Result | What it says |
|---|---|---|---|
| Estimated capacity | 100 requests/s/replica × 8 replicas | 800 requests/s | Assumes replicas scale near-linearly for this mix |
| Planning utilization | 820 requests/s ÷ 800 requests/s | 1.025 = 102.5% | Offered accepted work is above the estimate |
| Signed headroom | 800 − 820 requests/s | −20 requests/s | A 20 requests/s deficit in this model |
| Queue change rate | 820 arrivals/s − 760 completions/s | +60 requests/s | Accepted work accumulates if rates persist |
Over five minutes, the simple queue balance gives 60 requests/s × 300 s = 18,000 additional queued requests. This is a projection from constant rates, not a forecast: queues
can hit limits, shed work, trigger timeouts, or cause retries. The arithmetic is a useful consistency
check against the measured queue slope.
Eight times one replica is a hypothesis about scaling, too.
Multiplying 100 requests/s/replica × 8 replicas assumes those replicas contribute
independent capacity. Shared databases, locks, caches, network limits, or load imbalance can
break that assumption. A benchmark that saturates one replica against an isolated fake dependency
may measure a different system from production.
Capacity is also tied to a workload and a service objective. A tiny cached read and a large report-generation request each count as one request, but they do not consume equal work. When the mix changes, requests/s can hide a demand shift. Segment by route or use a workload-specific cost unit, and preserve p95, error rate, queue age, and rejected work alongside the aggregate rate.
Averages hide bursts and uneven distribution. A five-minute average below capacity can still miss a short spike or one hot replica. Sample shorter windows, inspect per-replica histograms, and compare the queue’s actual slope. A capacity number without its load shape and time window can create false confidence.
Adjust demand and replicas, then check whether the story still fits.
Capacity is per replica for the workload and objective described above; linear scaling is an assumption.
The queue estimate assumes accepted arrivals and completions share a boundary and that drops/cancellations are accounted for separately. Zero capacity has no finite utilization ratio.
Try increasing demand to 900 requests/s while holding eight replicas: the planning ratio rises to 112.5%, and estimated headroom falls to −100 requests/s. Now double replicas. The lab shows a lower aggregate ratio only because it assumes each additional replica brings another 100 requests/s of capacity. In a real system, verify that assumption against shared dependencies and a representative load test.
Keep the three calculations explicit.
The helpers return the planning utilization, signed headroom, and queue rate change. They validate finite, non-negative rates and refuse a zero capacity denominator. They cannot confirm that two callers measured the same workload, service objective, interval, or boundary; those assumptions belong in the metric contract and benchmark notes.
The examples retain request-rate units in names and label the ratio as a planning measure.
export type CapacityAssessment = {
utilization: number;
headroomPerSecond: number;
backlogChangePerSecond: number;
status: 'within-measured-capacity' | 'at-capacity' | 'over-capacity';
};
/** Compare offered demand and completed work with capacity for the same workload. */
export function assessCapacity(
offeredRequestsPerSecond: number,
completedRequestsPerSecond: number,
sustainableRequestsPerSecond: number
): CapacityAssessment {
for (const [name, value] of [
['offeredRequestsPerSecond', offeredRequestsPerSecond],
['completedRequestsPerSecond', completedRequestsPerSecond],
['sustainableRequestsPerSecond', sustainableRequestsPerSecond]
] as const) {
if (!Number.isFinite(value) || value < 0) {
throw new Error(`${name} must be finite and non-negative`);
}
}
if (sustainableRequestsPerSecond === 0) {
throw new Error('sustainableRequestsPerSecond must be positive');
}
const utilization = offeredRequestsPerSecond / sustainableRequestsPerSecond;
return {
utilization,
headroomPerSecond: sustainableRequestsPerSecond - offeredRequestsPerSecond,
backlogChangePerSecond: offeredRequestsPerSecond - completedRequestsPerSecond,
status:
utilization > 1
? 'over-capacity'
: utilization === 1
? 'at-capacity'
: 'within-measured-capacity'
};
}
package mathpractice
import (
"errors"
"math"
)
type CapacityAssessment struct {
Utilization float64
HeadroomPerSecond float64
BacklogChangePerSecond float64
Status string
}
// AssessCapacity compares offered demand and completed work with capacity
// measured for the same workload and service objective.
func AssessCapacity(offeredRequestsPerSecond, completedRequestsPerSecond, sustainableRequestsPerSecond float64) (CapacityAssessment, error) {
values := []struct {
name string
value float64
}{
{"offeredRequestsPerSecond", offeredRequestsPerSecond},
{"completedRequestsPerSecond", completedRequestsPerSecond},
{"sustainableRequestsPerSecond", sustainableRequestsPerSecond},
}
for _, input := range values {
if math.IsNaN(input.value) || math.IsInf(input.value, 0) || input.value < 0 {
return CapacityAssessment{}, errors.New(input.name + " must be finite and non-negative")
}
}
if sustainableRequestsPerSecond == 0 {
return CapacityAssessment{}, errors.New("sustainableRequestsPerSecond must be positive")
}
utilization := offeredRequestsPerSecond / sustainableRequestsPerSecond
status := "within-measured-capacity"
if utilization > 1 {
status = "over-capacity"
} else if utilization == 1 {
status = "at-capacity"
}
return CapacityAssessment{
Utilization: utilization,
HeadroomPerSecond: sustainableRequestsPerSecond - offeredRequestsPerSecond,
BacklogChangePerSecond: offeredRequestsPerSecond - completedRequestsPerSecond,
Status: status,
}, nil
}
A ratio can locate a mismatch; evidence locates the constraint.
Keep the model concise: U = D/C compares offered demand to measured capacity, H = C−D is the signed room left, and queue change ≈ arrivals−completions estimates accumulation over a short window. Each equation answers a different question, and
none identifies the bottleneck without measurements of the system.
For queue accounting, see John D. C. Little, “A Proof for the Queuing Formula: L = λW”, Operations Research 9(3), 1961. Little’s Law relates long-run averages in a stable system; the queue-slope calculation here is a separate, short-window balance.