Feature flag service.
Explain every decision.
Let a team release a feature gradually, see who receives it, and roll back a bad change. Build the small client library that keeps the application making decisions when the configuration service disappears.
The switch is the easy part. The depth is in stable rollout groups, versioned configuration, cached decisions, and knowing which application instances actually received a change.
A small product with a substantial boundary
A dashboard, a client library, and an app that uses it.
The first product loop
- Create a boolean flag in a workspace and environment.
- Preview it for a few synthetic users, then publish a 10% rollout.
- Run the demo application as two instances, each using your server-side TypeScript SDK.
- Increase the rollout to 25% and inspect the users who keep or gain access.
- Break configuration delivery to one instance, publish a rollback, and explain both instances’ behavior.
Keep the first version focused
Use one application codebase, a relational database, a worker, and a small TypeScript package. Start with boolean flags, explicit targeting keys, percentage rollouts, and a history of published snapshots.
The SDK runs in the demo application’s server. The browser receives evaluated results through that application; it never gets an administrative credential or a private targeting list.
Leave multivariate experiments, automatic statistical rollout decisions, and multiple SDK languages for later. This page is a project brief and drill planner; the service and SDK are what you would build.
Make the evaluator a contract
One snapshot, one identity, one explainable result.
- 01 / Author
Edit a draft
Validate rules and permissions. Preview against named synthetic users.
- 02 / Publish
Commit a revision
Atomically store an immutable snapshot, audit record, and pending notification work.
- 03 / Distribute
Refresh the SDK
Fetch and validate the full snapshot in the background, then swap it atomically.
- 04 / Evaluate
Decide locally
Return a value and reason from the installed snapshot without a network request.
The management service owns configuration changes. The SDK owns local evaluation and refresh. An evaluation should stay fast even while configuration delivery is failing.
Define the rule order before writing the controls
For this brief, a usable snapshot evaluates a boolean flag in this order: a disabled flag returns false; an explicit targeting-key override returns its configured value; otherwise the percentage rule returns whether the user’s bucket is below the threshold.
A disabled flag needs no targeting key. For an enabled flag, require a stable, nonempty targeting key; if it is missing, return the caller’s supplied default with an explanatory reason. An unknown flag, unsupported data, or an unavailable usable snapshot also returns the default. A false result and a failed evaluation are different outcomes.
Include the flag key, value, reason, snapshot revision, freshness, and relevant rule or bucket in the debug result. Capture one snapshot reference for a group of evaluations within the same application request so a refresh cannot mix revisions halfway through.
Make percentage rollouts deterministic
Give each user a stable bucket from 0 to 9,999. Specify the hash algorithm, byte encoding, field boundaries, and inputs: environment, flag key, a stable rollout salt, and targeting key. Store the salt with the flag. Keep publication revision and percentage out of those inputs.
A 10% rollout includes buckets below 1,000; 25% includes buckets below 2,500. With the same identities, salt, and targeting rules, the first cohort remains inside the second. Restarting an SDK does not reshuffle users.
Check fixed test vectors, empty or missing keys, and the 0% and 100% boundaries. A threshold does not guarantee an exact percentage in a small fixture. Unleash’s stickiness documentation is a useful concrete reference for stable assignments; your project still needs its own explicit contract.
Decide what stale configuration is allowed to do
Initialize asynchronously. Until a validated snapshot is ready, return the caller’s default. During an outage, a warm SDK may serve its last validated snapshot only within a declared maximum age; after that it returns the default with an expired reason.
Track freshness from the last successful authoritative refresh. Failed fetches and replayed update notices do not renew it. Begin with an in-memory cache: restarting the process is a cold start. Persisting an SDK cache is a later extension that must preserve age and handle clock changes.
For the demo feature, choose false as the caller’s fallback and show the age limit in the UI. That is a decision about this feature. Immediate global shutdown cannot be promised to disconnected instances that are still allowed to use cached rules.
Version updates and rollbacks in the same direction
Publish full, immutable snapshots with increasing revisions within an environment. Reject incomplete or incompatible candidates before installation. Ignore old updates, and treat duplicate notices as a reason to check current state rather than proof of freshness.
Rollback copies a previous snapshot’s rules into a new revision. Rolling revision 42 back to the rules from 41 produces revision 43. Existing SDK ordering rules can then adopt the rollback normally.
A schedule carries the revision its author expected. If another publication or rollback happens before execution, mark the schedule conflicted and require review. Do not silently rebase old scheduled intent onto the latest configuration.
Keep flag decisions separate from permissions
A flag controls which product behavior is offered. Application authorization still decides who may perform the underlying operation. A hidden button, a cached flag, or a client-supplied targeting key is not an access check.
Scope SDK credentials to the environment they can read, and enforce dashboard publishing permissions on the server. Explain that revoking a credential stops future fetches; it does not remotely erase configuration already cached by an offline process.
The portfolio moment
Two instances. One rollback. Different evidence.
Use this sequence as a demo script. These are proposed outcomes under the cache policy above, not measurements from a running service.
| Checkpoint | Published revision | Instance A | Instance B |
|---|---|---|---|
| Both connected | 41 | Evaluates 41 | Evaluates 41 |
| Increase rollout | 42 | Adopts 42 | Uses cached 41 within age limit |
| Roll back | 43, with rules from 41 | Adopts 43 | Still on cached 41; adoption unknown centrally |
| B’s cache expires | 43 | Evaluates 43 | Returns default false, reason expired |
| B reconnects | 43 | Evaluates 43 | Validates and adopts 43 |
A matching value does not establish a matching version. B can coincidentally serve the rollback’s old rules while having no knowledge that a rollback happened. Its local debug view can show the cache; the central dashboard needs fresh evidence before claiming adoption.
The engineering bar in this product
Turn each capability into a repeatable check.
01 / Authentication & credential lifecycle
Build
Use established session authentication for the dashboard. Give each demo application instance a read-only, environment-scoped SDK credential. Keep it on the application server.
Demonstrate
Expire an admin session and rotate an SDK credential. Further fetches with the old credential fail; document how long a disconnected instance can still use previously fetched data.
02 / Authorization
Build
Separate viewers, editors, and publishers. Enforce workspace and environment permissions on every mutation, snapshot read, and audit-history request.
Demonstrate
An editor cannot publish to production, and a credential for staging cannot fetch production configuration. Feature visibility never substitutes for application authorization.
03 / Persistence & atomic publication
Build
Persist drafts, immutable environment snapshots, a current-version pointer, audit entries, and publish work. Commit the new version, pointer, audit entry, and notification work together.
Demonstrate
Restart after a publish acknowledgement. The same version is current, its actor and change are recorded, and unfinished notification work can still be found.
04 / Background jobs
Build
Use durable jobs for scheduled publications and update notifications. Include an expected current revision, an operation identity, and expiring worker claims.
Demonstrate
Stop a worker after publication commits. Recovery does not publish twice. A scheduled change whose base revision is stale is marked conflicted instead of overwriting an emergency rollback.
05 / Retries & recovery
Build
Fetch snapshots with deadlines and capped backoff with jitter. Keep evaluation independent of network requests. Make notification retries bounded; periodic refresh remains a catch-up path.
Demonstrate
Return 500 or 429 from the configuration endpoint. Refresh attempts spread out, evaluations return promptly, and recovery installs the current snapshot without a synchronized retry burst.
06 / Idempotency & concurrent edits
Build
Require a workspace-scoped operation key for publication and rollback. Atomically bind the key to the request parameters and result. Compare the expected current revision before publishing.
Demonstrate
Retry the same publication after its response is lost: it returns the original version. Reusing its key for different content conflicts. Two edits against the same base cannot silently overwrite each other.
07 / Migrations & compatibility
Build
Version stored rules and the snapshot schema. Backfill existing flags before enforcing new constraints, and keep a declared compatibility window for older SDKs.
Demonstrate
Upgrade a populated database and fetch with the previous SDK version. An unsupported snapshot is rejected as a whole; it never becomes an accidentally empty set of flags.
08 / Rate limiting & bounded work
Build
Limit mutation and snapshot requests per workspace and credential. Bound notification queues, SDK refresh concurrency, and telemetry buffers; coalesce version notices when only the latest snapshot matters.
Demonstrate
Reconnect many demo instances with a bounded test. Show admitted and rejected requests, stable worker concurrency, and continued service for another workspace.
09 / Observability & decision explanations
Build
Expose the evaluated value, reason, snapshot revision, cache freshness, and rule or rollout bucket in a debug view. Correlate publications with SDK adoption; aggregate or sample evaluation telemetry.
Demonstrate
Explain why two instances served different values for the same user. Trace the publication and snapshot versions without putting user identifiers in metric labels or logging every evaluation.
10 / Monitoring & rollback visibility
Build
Track publication failures, refresh errors, active SDK version lag, stale/default evaluation rates, and telemetry gaps. Define alerts and a runbook for a declared workload.
Demonstrate
Disconnect one instance during rollback. The dashboard distinguishes publication committed, adoption observed, and adoption unknown. Show the alert clearing after confirmed recovery.
11 / Deterministic rollout
Build
Specify a stable hash and encoding of environment, flag, rollout salt, and targeting key. Keep revision and percentage out of the hash. Compare a fixed bucket with the configured threshold.
Demonstrate
Run fixed test vectors and the same synthetic users through both SDK instances. Increasing the percentage retains the earlier cohort when identity, salt, and targeting rules stay unchanged.
Build → break → explain
Make the configuration unreliable on purpose.
Build a local harness that can delay or reject refreshes, reorder notifications, interrupt workers, and return invalid snapshots. Show which instance is affected and keep the fault state visible. Use synthetic users and a bounded client count.
Availability & freshness
Configuration outage
- Inject this fault
- Block configuration fetches for one running instance, then start a second instance with no cached snapshot.
- Predict before you run it
- Do a warm instance and a cold instance have enough information to make the same decision?
- Behavior to build
- The warm instance may use its last validated snapshot within the declared age limit. Once that limit expires, it returns the caller’s default. The cold instance returns the default until initialization succeeds.
Evidence to collect
- Evaluation latency stays independent of the failing network request.
- The debug view distinguishes cached, expired, and not-ready results.
- A failed refresh never resets the cache freshness timer.
Recover & explain
Restore connectivity and adopt a validated current snapshot. Record the recovery time and which requests used a fallback.
This planner describes drills to implement in your project. Selecting a drill does not send traffic or inject a fault.
What the chaos controls should expose
Let the reviewer select an instance, a fault profile, and a bounded duration or request count. Display the current published revision beside each instance’s local revision, cache age, decision reason, and last observed report.
Provide a stop action that restores ordinary transport, plus a reset that recreates synthetic fixtures with a clear run identity. Keep the incident’s evidence available. A deliberately broken evaluator belongs in the test harness so cohort drift is observable without changing the normal SDK contract.
A buildable sequence
Grow from a pure decision to an operated service.
- 01
Make one decision explainable
Build a pure boolean evaluator, fixed targeting fixtures, and a demo app with two visible feature paths. Return value, reason, and revision.
Done when: The same snapshot and targeting key produce the same result. Off, explicit targeting, percentage boundaries, and missing context have documented behavior.
- 02
Publish with an audit trail
Add workspaces, environments, roles, drafts, immutable snapshots, atomic publication, and revision conflict checks.
Done when: Unauthorized publications fail. Concurrent edits cannot silently overwrite each other, and a repeated publish request returns its original result.
- 03
Give the SDK a lifecycle
Create a small server-side TypeScript client with initialization, background refresh, atomic snapshot swaps, defaults, freshness tracking, and shutdown.
Done when: Two instances stay responsive through outages, reject invalid or older snapshots, and explain cached versus default decisions.
- 04
Schedule, roll back, and operate
Add durable scheduled changes, revision-aware rollback, bounded retries, quotas, adoption telemetry, alerts, and a populated-data migration.
Done when: A stale scheduled increase cannot overwrite a rollback. You can distinguish published state from observed SDK adoption and explain missing evidence.
- 05
Package the failure story
Expose fault profiles in a local demo harness, with synthetic users, a bounded client generator, clear reset steps, and a short incident report.
Done when: A reviewer can reproduce a partitioned rollback and cohort-drift failure, inspect the evidence, and restore the system.
The portfolio handoff
Give the reviewer a decision they can investigate.
- Start with one user. Show the targeting key, stable bucket, value, and revision in both instances.
- Increase the rollout. Keep the existing cohort and explain who joins. Show the fixed fixture that verifies it.
- Partition one instance. Publish and roll back a change while its SDK cannot refresh.
- Inspect the disagreement. Trace the publication, installed revisions, age limit, and fallback. Distinguish local knowledge from central telemetry.
- Recover and qualify. Reconnect, verify adoption, and explain why this is neither instantaneous global rollback nor reversal of effects already performed.
Include the evaluator contract, test vectors, permission checks, a migration rehearsal, and a small incident report. Label propagation and latency measurements with the client count, workload, and environment. A debug display makes the reasoning inspectable; it does not replace verification.
Lessons behind this brief
Related lessons
The evaluator contract and the rollback drills build on these lessons.
- Feature flags & kill switches
Stable bucketing, safe stale-state behavior, and a switch on-call can use.
- Configuration as a boundary
Turn raw configuration into a validated, typed value before the application relies on it.
- Versioning & compatibility
Evolve a contract without surprising old consumers, such as an older SDK.
- Optimistic concurrency
Compare the expected revision so a stale schedule cannot overwrite a rollback.
Primary references
Use existing contracts to sharpen your own.
- OpenFeature: flag evaluation API
Defines caller defaults and detailed evaluation results. Use it to examine the shape of your client API; this brief does not claim OpenFeature compatibility.
- OpenFeature: evaluation context
Explains the targeting key and the context supplied to evaluation. Decide which identity your application owns and how it stays stable.
- Unleash: stickiness
A concrete implementation’s approach to consistent rollout assignments and grouping. Compare its choices with your fixture and hash contract.