The team needs a scale check before choosing an architecture.
For this illustrative planning case, the initial market estimate is 40,000 active accounts. Assume 20 audit events per account per day and an average serialized event of 1.5 KiB. The service should retain 30 days. The event rate will be bursty, and the full stored footprint will include indexes and replicas.
Those inputs are deliberately rounded assumptions, not telemetry. They are enough to discover whether the plan is closer to ten events a second or ten thousand, and whether raw retention looks like megabytes or terabytes. They are not enough to size a production cluster or promise a launch-day latency.
- Initial population
- 40,000 active accounts (planning assumption).
- Activity
- 20 events/account/day; event payload averages 1.5 KiB.
- Retention
- 30 days; storage overhead is not yet measured.
- Decision
- Is this a small queue and tens of GB, or an order of magnitude larger?
Keep each quantity attached to its unit and boundary.
An order of magnitude is a factor of ten. Saying “roughly ten events per second” instead of “9.259259 events/s” communicates the useful scale when account activity and payload size are estimates. Extra decimal places do not make guessed inputs more accurate.
Begin with a reference quantity you can defend. There are 86,400 seconds/day,
approximately 1,000 bytes/KB in decimal storage units, and 1,024 bytes/KiB in binary units. In this lesson, 1.5 KiB is 1,536 bytes; storage totals use
decimal GB (1 GB = 10⁹ bytes) so the stated units match common disk capacity
labels.
For throughput, calculate the average first, then apply an explicit peak factor. For storage, multiply event count by measured or assumed serialized bytes, retention, then a separate overhead factor. Do not hide index and replica growth inside “event size”: keeping the terms separate makes them easier to replace when measurements arrive.
Latency needs a different model. A rough budget can add stages that happen sequentially, but parallel work uses the longest branch, and adding stage p95 values does not in general produce an end-to-end p95. First draw the request path and state whether the reference is a measured median, a target, or a placeholder.
| Stage | Budget placeholder | Boundary to verify |
|---|---|---|
| Network and ingress | 12 ms | Client-to-service or internal hop? |
| Queue and application | 25 ms | Does this include queue wait? |
| Database operation | 35 ms | One query or the whole transaction? |
| Serialization and response | 20 ms | Does this include response transfer? |
| Simple serial sum | 92 ms | Leaves only 8 ms under a 100 ms target |
One base case is easy to calculate and easy to over-trust.
Vary the inputs that could change the answer most. The following cases are illustrative planning bounds, not percentiles or confidence intervals. Each row uses a consistent bundle of assumptions: account population, activity, payload size, peak factor, and storage overhead.
| Case | Inputs | Average → peak | Raw/day | 30-day stored estimate |
|---|---|---|---|---|
| Low | 20k accounts × 10/day × 0.75 KiB; 4× peak; 2× overhead | ≈ 2.3 → 9 events/s | ≈ 0.15 GB | ≈ 9 GB |
| Base | 40k × 20/day × 1.5 KiB; 8× peak; 3× overhead | ≈ 9.3 → 74 events/s | ≈ 1.23 GB | ≈ 111 GB |
| High | 80k × 40/day × 3 KiB; 15× peak; 5× overhead | ≈ 37 → 556 events/s | ≈ 9.83 GB | ≈ 1.47 TB |
The high case is about thirteen times the base stored footprint. That gap is a useful finding: account activity, average encoded size, and expansion are worth measuring before a purchase or partitioning decision. The bounds are deliberately not presented as “likely” probabilities; we have not measured a distribution for any input.
A quick independent check catches arithmetic slips. In the base case, 800,000 events/day at roughly 1.5 kB each is a little over 1 GB/day. Thirty days is around 30–40 GB raw. A 3× allowance puts the expanded footprint near 100 GB, not 10 GB or 1 TB.
Move an input and see which order of magnitude changes.
Choose a starting case, then edit one assumption at a time. The builder uses decimal GB for storage, 1 KiB = 1,024 bytes for payloads, and a flat storage expansion factor. It keeps the calculation visible; it does not model compression, compaction, uneven tenants, or burst timing.
Audit pipeline sizing
40,000 accounts × 20 events/account/day ÷ 86,400 s/day = 9.3 average events/s. Multiply by 8 for the peak planning rate.
The arithmetic is exact; the fit of the model is uncertain.
The arithmetic assumes the average event rate is a good daily summary and a single multiplier captures peaks. Real usage can cluster by hour, tenant, or release. A peak factor of eight says nothing about whether bursts last one second or two hours, even though those cases stress a queue very differently.
The storage factor is also not a law. Compression, indexes, replication, tombstones, compaction, backups, and retention semantics can each change the footprint. “1.5 KiB event” could mean the JSON body before transport, the compressed wire size, or the actual database row. Choose one boundary, then measure representative records at the next boundary.
Latency is especially easy to overclaim. The illustrative stage budgets sum to 92 ms only if all stages are sequential and use comparable observations. Parallel branches overlap; queueing adds load-dependent delay; and component percentiles do not simply add to the same end-to-end percentile. Instrument one request with spans, record the full distribution, and compare the same request class against its target.
Make the assumptions explicit inputs.
The helpers calculate events/day, average and peak events/second, raw daily bytes, and a retained storage estimate. The examples validate non-finite and negative inputs, and use named fields so units are visible at the call site. The estimate remains a model: code cannot make unmeasured activity or storage expansion factual.
Both examples preserve the shared inputs and convert to display units only after the byte arithmetic.
export type EstimateInput = {
accounts: number;
eventsPerAccountPerDay: number;
bytesPerEvent: number;
retentionDays: number;
peakFactor: number;
storageFactor: number;
};
export type Estimate = {
eventsPerDay: number;
averageEventsPerSecond: number;
peakEventsPerSecond: number;
rawBytesPerDay: number;
retainedRawBytes: number;
estimatedStoredBytes: number;
};
export function estimatePipeline(input: EstimateInput): Estimate {
const values = Object.values(input);
if (!values.every(Number.isFinite) || values.some((value) => value < 0)) {
throw new Error('Inputs must be finite, non-negative numbers.');
}
if (input.accounts === 0 || input.bytesPerEvent === 0 || input.retentionDays === 0) {
return {
eventsPerDay: 0,
averageEventsPerSecond: 0,
peakEventsPerSecond: 0,
rawBytesPerDay: 0,
retainedRawBytes: 0,
estimatedStoredBytes: 0
};
}
const eventsPerDay = input.accounts * input.eventsPerAccountPerDay;
const averageEventsPerSecond = eventsPerDay / 86_400;
const rawBytesPerDay = eventsPerDay * input.bytesPerEvent;
const retainedRawBytes = rawBytesPerDay * input.retentionDays;
return {
eventsPerDay,
averageEventsPerSecond,
peakEventsPerSecond: averageEventsPerSecond * input.peakFactor,
rawBytesPerDay,
retainedRawBytes,
estimatedStoredBytes: retainedRawBytes * input.storageFactor
};
}
package main
import (
"errors"
"fmt"
"math"
)
const secondsPerDay = 86_400.0
type EstimateInput struct {
Accounts float64
EventsPerAccountPerDay float64
BytesPerEvent float64
RetentionDays float64
PeakFactor float64
StorageFactor float64
}
type Estimate struct {
EventsPerDay float64
AverageEventsPerSecond float64
PeakEventsPerSecond float64
RawBytesPerDay float64
RetainedRawBytes float64
EstimatedStoredBytes float64
}
func EstimatePipeline(input EstimateInput) (Estimate, error) {
values := []float64{
input.Accounts,
input.EventsPerAccountPerDay,
input.BytesPerEvent,
input.RetentionDays,
input.PeakFactor,
input.StorageFactor,
}
for _, value := range values {
if math.IsNaN(value) || math.IsInf(value, 0) || value < 0 {
return Estimate{}, errors.New("inputs must be finite, non-negative numbers")
}
}
if input.Accounts == 0 || input.BytesPerEvent == 0 || input.RetentionDays == 0 {
return Estimate{}, nil
}
eventsPerDay := input.Accounts * input.EventsPerAccountPerDay
averageEventsPerSecond := eventsPerDay / secondsPerDay
rawBytesPerDay := eventsPerDay * input.BytesPerEvent
retainedRawBytes := rawBytesPerDay * input.RetentionDays
return Estimate{
EventsPerDay: eventsPerDay,
AverageEventsPerSecond: averageEventsPerSecond,
PeakEventsPerSecond: averageEventsPerSecond * input.PeakFactor,
RawBytesPerDay: rawBytesPerDay,
RetainedRawBytes: retainedRawBytes,
EstimatedStoredBytes: retainedRawBytes * input.StorageFactor,
}, nil
}
func main() {
input := EstimateInput{
Accounts: 40_000, EventsPerAccountPerDay: 20, BytesPerEvent: 1_536,
RetentionDays: 30, PeakFactor: 8, StorageFactor: 3,
}
estimate, err := EstimatePipeline(input)
if err != nil {
panic(err)
}
fmt.Printf("events/day: %.0f\naverage events/s: %.1f\npeak events/s: %.0f\n", estimate.EventsPerDay, estimate.AverageEventsPerSecond, estimate.PeakEventsPerSecond)
fmt.Printf("raw GB/day: %.2f\nestimated stored GB: %.1f\n", estimate.RawBytesPerDay/1e9, estimate.EstimatedStoredBytes/1e9)
}
A useful estimate ends with a measurement plan.
| Assumption | Measure next | Useful boundary |
|---|---|---|
| Accounts × events/account/day | Events per tenant and per hour | Active accounts, timezone, and event type |
| 1.5 KiB/event | Serialized and persisted bytes/event | Before compression and after database encoding |
| 8× average peak | Short-window arrival-rate distribution | Seconds, minutes, and sustained busy intervals |
| 3× storage expansion | Physical bytes per logical retained byte | Indexes, replication, backups, and compaction state |
| 92 ms stage budget | End-to-end traces and latency percentiles | Same request class, target, and full path |
The habit transfers to any sizing question: write down the decision, name the population and boundary, use a reference number with units, calculate a base case, vary the assumptions that matter, and find an independent sanity check. Then ask which observation would most reduce the uncertainty before you act.