← Math in Practice
Concept Math in engineering decisions

Back-of-envelope estimation

Use rough arithmetic to ask a better question before the system exists.

A product team wants to ship an audit-event pipeline before a customer launch. The request sounds simple: “Can we handle the traffic, keep 30 days, and stay under 100 ms?” There is no production traffic to inspect yet. The useful first answer is not a confident yes or no; it is a transparent range that shows which unknowns could change the decision.

The judgment to keep

Estimate with numbers you can explain, round to the precision those numbers deserve, and mark which assumptions need measurement before the estimate becomes an engineering commitment.

TypeScriptGo Order of magnitude · throughput · storage · latency budget · measurement
01 / Read the launch request

The team needs a scale check before choosing an architecture.

For this illustrative planning case, the initial market estimate is 40,000 active accounts. Assume 20 audit events per account per day and an average serialized event of 1.5 KiB. The service should retain 30 days. The event rate will be bursty, and the full stored footprint will include indexes and replicas.

Those inputs are deliberately rounded assumptions, not telemetry. They are enough to discover whether the plan is closer to ten events a second or ten thousand, and whether raw retention looks like megabytes or terabytes. They are not enough to size a production cluster or promise a launch-day latency.

Case file / Audit pipeline Estimate demand and retained footprint, then identify the biggest uncertainty.
Initial population
40,000 active accounts (planning assumption).
Activity
20 events/account/day; event payload averages 1.5 KiB.
Retention
30 days; storage overhead is not yet measured.
Decision
Is this a small queue and tens of GB, or an order of magnitude larger?
First conclusion: with these assumptions, average throughput is single-digit events/s, peak planning demand is tens per second, and raw 30-day storage is tens of GB.
02 / Choose useful reference numbers

Keep each quantity attached to its unit and boundary.

An order of magnitude is a factor of ten. Saying “roughly ten events per second” instead of “9.259259 events/s” communicates the useful scale when account activity and payload size are estimates. Extra decimal places do not make guessed inputs more accurate.

Begin with a reference quantity you can defend. There are 86,400 seconds/day, approximately 1,000 bytes/KB in decimal storage units, and 1,024 bytes/KiB in binary units. In this lesson, 1.5 KiB is 1,536 bytes; storage totals use decimal GB (1 GB = 10⁹ bytes) so the stated units match common disk capacity labels.

For throughput, calculate the average first, then apply an explicit peak factor. For storage, multiply event count by measured or assumed serialized bytes, retention, then a separate overhead factor. Do not hide index and replica growth inside “event size”: keeping the terms separate makes them easier to replace when measurements arrive.

Latency needs a different model. A rough budget can add stages that happen sequentially, but parallel work uses the longest branch, and adding stage p95 values does not in general produce an end-to-end p95. First draw the request path and state whether the reference is a measured median, a target, or a placeholder.

Average event rateaccounts × events/account/day ÷ 86,400 s/day
Peak planning rateaverage rate × peak factor
Retained storageevents/day × bytes/event × days × overhead
Illustrative serial-path budget; these are planning placeholders, not observed service data
StageBudget placeholderBoundary to verify
Network and ingress12 msClient-to-service or internal hop?
Queue and application25 msDoes this include queue wait?
Database operation35 msOne query or the whole transaction?
Serialization and response20 msDoes this include response transfer?
Simple serial sum92 msLeaves only 8 ms under a 100 ms target
03 / Build a range

One base case is easy to calculate and easy to over-trust.

Vary the inputs that could change the answer most. The following cases are illustrative planning bounds, not percentiles or confidence intervals. Each row uses a consistent bundle of assumptions: account population, activity, payload size, peak factor, and storage overhead.

Three plausible planning cases · decimal GB for storage totals
CaseInputsAverage → peakRaw/day30-day stored estimate
Low20k accounts × 10/day × 0.75 KiB; 4× peak; 2× overhead≈ 2.3 → 9 events/s≈ 0.15 GB≈ 9 GB
Base40k × 20/day × 1.5 KiB; 8× peak; 3× overhead≈ 9.3 → 74 events/s≈ 1.23 GB≈ 111 GB
High80k × 40/day × 3 KiB; 15× peak; 5× overhead≈ 37 → 556 events/s≈ 9.83 GB≈ 1.47 TB

The high case is about thirteen times the base stored footprint. That gap is a useful finding: account activity, average encoded size, and expansion are worth measuring before a purchase or partitioning decision. The bounds are deliberately not presented as “likely” probabilities; we have not measured a distribution for any input.

A quick independent check catches arithmetic slips. In the base case, 800,000 events/day at roughly 1.5 kB each is a little over 1 GB/day. Thirty days is around 30–40 GB raw. A 3× allowance puts the expanded footprint near 100 GB, not 10 GB or 1 TB.

04 / Change the assumptions

Move an input and see which order of magnitude changes.

Choose a starting case, then edit one assumption at a time. The builder uses decimal GB for storage, 1 KiB = 1,024 bytes for payloads, and a flat storage expansion factor. It keeps the calculation visible; it does not model compression, compaction, uneven tenants, or burst timing.

Interactive estimate / illustrative inputs

Audit pipeline sizing

Events per day800,000
Average rate9.3 events/s
Peak planning rate74.1 events/s
Raw storage/day1.23 GB
Expanded retained estimate110.59 GB

40,000 accounts × 20 events/account/day ÷ 86,400 s/day = 9.3 average events/s. Multiply by 8 for the peak planning rate.

05 / Challenge the estimate

The arithmetic is exact; the fit of the model is uncertain.

The arithmetic assumes the average event rate is a good daily summary and a single multiplier captures peaks. Real usage can cluster by hour, tenant, or release. A peak factor of eight says nothing about whether bursts last one second or two hours, even though those cases stress a queue very differently.

The storage factor is also not a law. Compression, indexes, replication, tombstones, compaction, backups, and retention semantics can each change the footprint. “1.5 KiB event” could mean the JSON body before transport, the compressed wire size, or the actual database row. Choose one boundary, then measure representative records at the next boundary.

Latency is especially easy to overclaim. The illustrative stage budgets sum to 92 ms only if all stages are sequential and use comparable observations. Parallel branches overlap; queueing adds load-dependent delay; and component percentiles do not simply add to the same end-to-end percentile. Instrument one request with spans, record the full distribution, and compare the same request class against its target.

06 / Practice in code

Make the assumptions explicit inputs.

The helpers calculate events/day, average and peak events/second, raw daily bytes, and a retained storage estimate. The examples validate non-finite and negative inputs, and use named fields so units are visible at the call site. The estimate remains a model: code cannot make unmeasured activity or storage expansion factual.

Compare the same estimate in TypeScript and Go.

Both examples preserve the shared inputs and convert to display units only after the byte arithmetic.

TypeScriptEvent pipeline sizing · explicit assumptions
estimate.ts · event pipeline sizing
export type EstimateInput = {
	accounts: number;
	eventsPerAccountPerDay: number;
	bytesPerEvent: number;
	retentionDays: number;
	peakFactor: number;
	storageFactor: number;
};

export type Estimate = {
	eventsPerDay: number;
	averageEventsPerSecond: number;
	peakEventsPerSecond: number;
	rawBytesPerDay: number;
	retainedRawBytes: number;
	estimatedStoredBytes: number;
};

export function estimatePipeline(input: EstimateInput): Estimate {
	const values = Object.values(input);
	if (!values.every(Number.isFinite) || values.some((value) => value < 0)) {
		throw new Error('Inputs must be finite, non-negative numbers.');
	}
	if (input.accounts === 0 || input.bytesPerEvent === 0 || input.retentionDays === 0) {
		return {
			eventsPerDay: 0,
			averageEventsPerSecond: 0,
			peakEventsPerSecond: 0,
			rawBytesPerDay: 0,
			retainedRawBytes: 0,
			estimatedStoredBytes: 0
		};
	}

	const eventsPerDay = input.accounts * input.eventsPerAccountPerDay;
	const averageEventsPerSecond = eventsPerDay / 86_400;
	const rawBytesPerDay = eventsPerDay * input.bytesPerEvent;
	const retainedRawBytes = rawBytesPerDay * input.retentionDays;

	return {
		eventsPerDay,
		averageEventsPerSecond,
		peakEventsPerSecond: averageEventsPerSecond * input.peakFactor,
		rawBytesPerDay,
		retainedRawBytes,
		estimatedStoredBytes: retainedRawBytes * input.storageFactor
	};
}
GoEvent pipeline sizing · explicit assumptions
estimate.go · event pipeline sizing
package main

import (
	"errors"
	"fmt"
	"math"
)

const secondsPerDay = 86_400.0

type EstimateInput struct {
	Accounts               float64
	EventsPerAccountPerDay float64
	BytesPerEvent          float64
	RetentionDays          float64
	PeakFactor             float64
	StorageFactor          float64
}

type Estimate struct {
	EventsPerDay           float64
	AverageEventsPerSecond float64
	PeakEventsPerSecond    float64
	RawBytesPerDay         float64
	RetainedRawBytes       float64
	EstimatedStoredBytes   float64
}

func EstimatePipeline(input EstimateInput) (Estimate, error) {
	values := []float64{
		input.Accounts,
		input.EventsPerAccountPerDay,
		input.BytesPerEvent,
		input.RetentionDays,
		input.PeakFactor,
		input.StorageFactor,
	}
	for _, value := range values {
		if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 {
			return Estimate{}, errors.New("inputs must be finite, non-negative numbers")
		}
	}
	if input.Accounts == 0 || input.BytesPerEvent == 0 || input.RetentionDays == 0 {
		return Estimate{}, nil
	}

	eventsPerDay := input.Accounts * input.EventsPerAccountPerDay
	averageEventsPerSecond := eventsPerDay / secondsPerDay
	rawBytesPerDay := eventsPerDay * input.BytesPerEvent
	retainedRawBytes := rawBytesPerDay * input.RetentionDays

	return Estimate{
		EventsPerDay:           eventsPerDay,
		AverageEventsPerSecond: averageEventsPerSecond,
		PeakEventsPerSecond:    averageEventsPerSecond * input.PeakFactor,
		RawBytesPerDay:         rawBytesPerDay,
		RetainedRawBytes:       retainedRawBytes,
		EstimatedStoredBytes:   retainedRawBytes * input.StorageFactor,
	}, nil
}

func main() {
	input := EstimateInput{
		Accounts: 40_000, EventsPerAccountPerDay: 20, BytesPerEvent: 1_536,
		RetentionDays: 30, PeakFactor: 8, StorageFactor: 3,
	}
	estimate, err := EstimatePipeline(input)
	if err != nil {
		panic(err)
	}
	fmt.Printf("events/day: %.0f\naverage events/s: %.1f\npeak events/s: %.0f\n", estimate.EventsPerDay, estimate.AverageEventsPerSecond, estimate.PeakEventsPerSecond)
	fmt.Printf("raw GB/day: %.2f\nestimated stored GB: %.1f\n", estimate.RawBytesPerDay/1e9, estimate.EstimatedStoredBytes/1e9)
}
07 / Choose what to measure

A useful estimate ends with a measurement plan.

Turn the largest assumptions into observable quantities
AssumptionMeasure nextUseful boundary
Accounts × events/account/dayEvents per tenant and per hourActive accounts, timezone, and event type
1.5 KiB/eventSerialized and persisted bytes/eventBefore compression and after database encoding
8× average peakShort-window arrival-rate distributionSeconds, minutes, and sustained busy intervals
3× storage expansionPhysical bytes per logical retained byteIndexes, replication, backups, and compaction state
92 ms stage budgetEnd-to-end traces and latency percentilesSame request class, target, and full path

The habit transfers to any sizing question: write down the decision, name the population and boundary, use a reference number with units, calculate a base case, vary the assumptions that matter, and find an independent sanity check. Then ask which observation would most reduce the uncertainty before you act.