← Concepts & practices
Pattern Boundaries and contracts

Feature flags & kill switches

Change behavior at runtime, and give the decision a way to end.

A thumbnail service is rolling out a new pipeline to 10% of users. The rollout needs a stable assignment so a user does not switch paths on every request, a safe answer when the flag provider is stale, and a kill switch that on-call can use during an incident. The interesting part is not the boolean. It is who evaluates the decision, what evidence travels with it, and when the temporary branch disappears.

The judgment to keep

A feature flag is an owned runtime policy, not a global boolean. Evaluate it from explicit context, return a safe path and a reason, observe the decision, and attach an owner, expiry, and retirement plan. A kill switch is the operational half of that contract: it must be safe to use when everything else is under pressure.

TypeScriptGo One thumbnail rollout · four runtime states.
Start with the decision

A flag changes behavior over time, not just in code.

A normal conditional answers “what does this input mean?” A flag also answers “who gets this behavior now, based on which version of policy, and what happens when the source of that policy is unavailable?” Those extra questions are why a flag deserves a boundary.

For the thumbnail pipeline, newThumbnailPipeline = true is not enough. Ten percent of which users? Is the assignment stable? Is a 45-minute-old snapshot still trustworthy? Does an incident responder have a named control? Can support explain why one asset used the new path?

Make the runtime decision explicit before making it dynamic.

Read the tempting versionTypeScript · a boolean leaks its missing policy
flags.ts · caller-local boolean
// The unsafe path makes a runtime choice without a stable subject or an owner.
export type RawFlagSource = Readonly<{ newThumbnailPipeline?: boolean }>;

export function rawThumbnailPath(flags: RawFlagSource): ThumbnailPath {
	return flags.newThumbnailPipeline === true ? 'next' : 'legacy';
}

export function randomPerRequestAssignment(): boolean {
	return Math.random() < 0.1;
}

The boolean is easy to read, but it does not say who owns rollout percentage, whether the value is fresh, or how to stop the path during an incident. A random assignment is even less useful: the same user can receive different behavior on adjacent requests.

Choose the owner

The flag’s location determines who can explain its behavior.

There is a useful progression from a local conditional to a runtime policy. The goal is not to turn every if into a service call. It is to make the extra lifetime, targeting, and operational decisions visible when they are real.

Caller

Consumes a decision

It provides context and chooses what to render or execute for the returned path.

Evaluator

Owns policy semantics

It handles targeting, stable bucketing, freshness, defaults, reasons, and version facts.

Operations

Owns the emergency control

It can activate a safe-off switch, see its effect, and retire the temporary control.

Three homes for a runtime decision
DesignWhat it can answerWhat it misses
Inline booleanIs this path enabled in this caller?Stable assignment, freshness, owner, and incident history.
Central evaluatorWhich variant does this context receive, and why?A named emergency control unless the contract includes one.
Snapshot + kill switchCan this version safely choose a path right now?Nothing for free: it still needs expiry, telemetry, and tests.
Read the Go evaluatorGo · explicit snapshot and safe decision
boundary.go · owned evaluator
func StableBucket(userID string, key string) int {
	hash := fnv.New32a()
	_, _ = hash.Write([]byte(key + ":" + userID))
	return int(hash.Sum32() % 100)
}

func Evaluate(snapshot *FlagSnapshot, context FlagContext, now time.Time) Decision {
	if snapshot == nil {
		return Decision{Path: "legacy", Reason: "default-off"}
	}
	if snapshot.KillSwitches[killSwitchKey] {
		return Decision{Path: "legacy", Reason: "kill-switch", FlagVersion: snapshot.Version}
	}
	definition, ok := snapshot.Flags[flagKey]
	if !ok {
		return Decision{Path: "legacy", Reason: "missing-flag", FlagVersion: snapshot.Version}
	}
	age := now.Sub(snapshot.FetchedAt)
	if age < 0 || age > snapshot.MaxAge {
		return Decision{Path: "legacy", Reason: "stale-snapshot", FlagVersion: snapshot.Version}
	}
	if definition.RolloutPercent <= 0 {
		return Decision{Path: "legacy", Reason: "default-off", FlagVersion: snapshot.Version}
	}
	bucket := StableBucket(context.UserID, flagKey)
	if bucket < definition.RolloutPercent {
		return Decision{Path: "next", Reason: "rollout", FlagVersion: snapshot.Version, Bucket: &bucket}
	}
	return Decision{Path: "legacy", Reason: "outside-rollout", FlagVersion: snapshot.Version, Bucket: &bucket}
}

The evaluator accepts a snapshot and a clock instead of reading process-global state. A stale snapshot, missing flag, or active switch returns the legacy path with a reason that can be logged and tested.

Follow the flag

Change the state and watch the ownership move.

Try the same user and rollout story with three designs. Then switch to a stale provider and an active incident. The safe design does not merely return “off”: it tells you whether the path was off by rollout, default, staleness, or an operator’s switch.

Thumbnail rollout

Trace the same flag through a runtime scenario.

Runs a local policy model
The next path is allowed for this context.new-path
01Receive context and the current flag state
02Read a caller-local boolean
03Choose a stable variant for this user
04Return the path, reason, and telemetry facts together
Evaluation

if (flags.newThumbnailPipeline) — every caller makes its own decision.

Ownership

No clear owner: each caller can interpret the flag.

Telemetry

Path choice is hard to compare with flag version, user, or incident time.

Watch the lifetime

A flag without a single evaluator becomes an undocumented, permanent second configuration system.

The controls model evaluation; they do not call a flag service or persist a switch.
Test the lifetime

Ask what happens before, during, and after the rollout.

The dangerous cases are often temporal: assignment changes between requests, a provider stops refreshing, an operator needs to act, or a “temporary” flag survives three quarters. Choose the policy that keeps those moments legible.

How should a 10% rollout choose whether a user gets the new path?
The flag snapshot is older than its allowed max age. What should a risky thumbnail path do?
What makes a kill switch operationally real?
The rollout is complete and the new path is the default. What is the next step?
Feedback stays on this page; it is not saved.
Operate the switch

A production flag is a small control plane.

The example keeps the evaluator pure: snapshot in, decision out. The composition layer supplies refreshing, a clock, and telemetry. That separation makes the decision testable without a flag vendor and makes the runtime integration responsible for the things only runtime can know.

For each flag, record the key, purpose, owner, created date, expiry or review date, target context, default behavior, kill-switch behavior, dashboard, alert, and removal issue. Test both the new and legacy paths, plus missing, malformed, stale, and switched-off state. Verify the switch in a controlled exercise before the incident that makes it urgent.

A caller-local boolean and random assignment make the rollout hard to reason about.

TypeScriptReading
flags.ts
// The unsafe path makes a runtime choice without a stable subject or an owner.
export type RawFlagSource = Readonly<{ newThumbnailPipeline?: boolean }>;

export function rawThumbnailPath(flags: RawFlagSource): ThumbnailPath {
	return flags.newThumbnailPipeline === true ? 'next' : 'legacy';
}

export function randomPerRequestAssignment(): boolean {
	return Math.random() < 0.1;
}
GoAlongside
boundary.go
type RawFlagSource struct {
	NewThumbnailPipeline bool
}

func RawThumbnailPath(flags RawFlagSource) string {
	if flags.NewThumbnailPipeline {
		return "next"
	}
	return "legacy"
}
Assignment

Make it stable

Hash a stable subject and flag key; never make a canary a per-request lottery.

Safety

Fail deliberately

Missing or stale state chooses the safe path and carries a reason for operators.

Lifetime

Remove the branch

When evidence says “default on,” delete the flag, old path, dashboards, and runbook step.

Recognize it in UI code

Rendering code should receive a decision, not a flag service.

Build frontends?Keep runtime policy outside rendering code.

Where it already is in your components

A component that checks flags.newThumbnailPipeline before choosing what to render is already evaluating a flag. If every screen reads the flag store directly, every screen owns the rollout.

When you have to own it

When a flag decides what a page shows, own where it is evaluated. The textbook components receive a domain-facing thumbnail decision and render the chosen asset. The wild components read a public global directly, so storage shape and rollout vocabulary now belong to each screen. That coupling makes a provider migration or emergency policy change a UI search.

The UI receives a thumbnail decision from the application boundary and only renders it.

ReactAlready in your code
textbook.tsx · decision projection
type ThumbnailDecision = { path: 'legacy' | 'next'; reason: string };
type ThumbnailApp = { getThumbnailDecision(assetId: string): ThumbnailDecision };

export function Thumbnail({ app, assetId }: { app: ThumbnailApp; assetId: string }) {
	const decision = app.getThumbnailDecision(assetId);
	return (
		<figure data-path={decision.path}>
			<img
				src={decision.path === 'next' ? `/next/${assetId}.webp` : `/legacy/${assetId}.jpg`}
				alt=""
			/>
			<figcaption>
				{decision.path} · {decision.reason}
			</figcaption>
		</figure>
	);
}
Keep the flag honest

Most flag failures are ownership and lifetime failures.

01

Random per request

A user can see both paths in one session, making bugs and metrics impossible to reproduce.

02

Stale means “probably fine”

Do not widen a risky rollout because the last snapshot is convenient. Put max age and fallback in the contract.

03

Switch without an audit

A switch change needs an actor, timestamp, reason, notification, and a way to confirm propagation.

04

Flag debt

Every live flag adds states to the system. A missing removal date turns a release tool into permanent complexity.

Do not hide a business rule in a flagRuntime release control is not domain policy

“Show the new thumbnail pipeline to 10% of users” is release policy. “A premium account may generate 50 thumbnails per hour” is a business rule. The latter needs a domain or authorization owner and must not become an accidental flag that a release operator can change without the right controls.

Make the call

Use a flag when change needs a controlled runtime window.

A flag earns its place for a staged release, a canary, a temporary migration, an experiment with measurable exposure, or an operational off switch. Keep it close to the boundary that owns the decision, make its safe behavior explicit, and give it a review or removal date before it ships.

Do not add a flag to avoid choosing a permanent domain rule, to replace authorization, or to make a static deployment setting dynamic without a real operator. If the behavior is now simply the new default, the flag has completed its job.

Keep this questionAsk it before adding the next boolean.

Who evaluates this decision, what is the safe answer when its state is missing or stale, who can turn it off, and what evidence will let us delete it?

Take the idea with you

A flag is a boundary around a temporary decision.

Make the decision safe to change, observable while it lives, and easy to remove.

Connections to follow nextRelated lessons
Why
The new thumbnail pipeline needs a staged, reversible rollout and an off switch during an incident.
What
One evaluator returns a path and a reason from a stable bucket of the user and flag key.
Constraint
A user stays on one path across requests, and on-call can use the kill switch without a deploy.
Fallback
A missing flag, a stale snapshot, or the kill switch returns the legacy path, each with its own reason.
Reconsider when
The rollout reaches everyone with evidence; then delete the flag and the old branch.