← Projects with depth
Project brief / 02Preview

Feature flag service.
Explain every decision.

Let a team release a feature gradually, see who receives it, and roll back a bad change. Build the small client library that keeps the application making decisions when the configuration service disappears.

The switch is the easy part. The depth is in stable rollout groups, versioned configuration, cached decisions, and knowing which application instances actually received a change.

A small product with a substantial boundary

A dashboard, a client library, and an app that uses it.

The first product loop

  1. Create a boolean flag in a workspace and environment.
  2. Preview it for a few synthetic users, then publish a 10% rollout.
  3. Run the demo application as two instances, each using your server-side TypeScript SDK.
  4. Increase the rollout to 25% and inspect the users who keep or gain access.
  5. Break configuration delivery to one instance, publish a rollback, and explain both instances’ behavior.

Keep the first version focused

Use one application codebase, a relational database, a worker, and a small TypeScript package. Start with boolean flags, explicit targeting keys, percentage rollouts, and a history of published snapshots.

The SDK runs in the demo application’s server. The browser receives evaluated results through that application; it never gets an administrative credential or a private targeting list.

Leave multivariate experiments, automatic statistical rollout decisions, and multiple SDK languages for later. This page is a project brief and drill planner; the service and SDK are what you would build.

Make the evaluator a contract

One snapshot, one identity, one explainable result.

  1. 01 / Author

    Edit a draft

    Validate rules and permissions. Preview against named synthetic users.

  2. 02 / Publish

    Commit a revision

    Atomically store an immutable snapshot, audit record, and pending notification work.

  3. 03 / Distribute

    Refresh the SDK

    Fetch and validate the full snapshot in the background, then swap it atomically.

  4. 04 / Evaluate

    Decide locally

    Return a value and reason from the installed snapshot without a network request.

The management service owns configuration changes. The SDK owns local evaluation and refresh. An evaluation should stay fast even while configuration delivery is failing.

Define the rule order before writing the controls

For this brief, a usable snapshot evaluates a boolean flag in this order: a disabled flag returns false; an explicit targeting-key override returns its configured value; otherwise the percentage rule returns whether the user’s bucket is below the threshold.

A disabled flag needs no targeting key. For an enabled flag, require a stable, nonempty targeting key; if it is missing, return the caller’s supplied default with an explanatory reason. An unknown flag, unsupported data, or an unavailable usable snapshot also returns the default. A false result and a failed evaluation are different outcomes.

Include the flag key, value, reason, snapshot revision, freshness, and relevant rule or bucket in the debug result. Capture one snapshot reference for a group of evaluations within the same application request so a refresh cannot mix revisions halfway through.

Make percentage rollouts deterministic

Give each user a stable bucket from 0 to 9,999. Specify the hash algorithm, byte encoding, field boundaries, and inputs: environment, flag key, a stable rollout salt, and targeting key. Store the salt with the flag. Keep publication revision and percentage out of those inputs.

A 10% rollout includes buckets below 1,000; 25% includes buckets below 2,500. With the same identities, salt, and targeting rules, the first cohort remains inside the second. Restarting an SDK does not reshuffle users.

Check fixed test vectors, empty or missing keys, and the 0% and 100% boundaries. A threshold does not guarantee an exact percentage in a small fixture. Unleash’s stickiness documentation is a useful concrete reference for stable assignments; your project still needs its own explicit contract.

Decide what stale configuration is allowed to do

Initialize asynchronously. Until a validated snapshot is ready, return the caller’s default. During an outage, a warm SDK may serve its last validated snapshot only within a declared maximum age; after that it returns the default with an expired reason.

Track freshness from the last successful authoritative refresh. Failed fetches and replayed update notices do not renew it. Begin with an in-memory cache: restarting the process is a cold start. Persisting an SDK cache is a later extension that must preserve age and handle clock changes.

For the demo feature, choose false as the caller’s fallback and show the age limit in the UI. That is a decision about this feature. Immediate global shutdown cannot be promised to disconnected instances that are still allowed to use cached rules.

Version updates and rollbacks in the same direction

Publish full, immutable snapshots with increasing revisions within an environment. Reject incomplete or incompatible candidates before installation. Ignore old updates, and treat duplicate notices as a reason to check current state rather than proof of freshness.

Rollback copies a previous snapshot’s rules into a new revision. Rolling revision 42 back to the rules from 41 produces revision 43. Existing SDK ordering rules can then adopt the rollback normally.

A schedule carries the revision its author expected. If another publication or rollback happens before execution, mark the schedule conflicted and require review. Do not silently rebase old scheduled intent onto the latest configuration.

Keep flag decisions separate from permissions

A flag controls which product behavior is offered. Application authorization still decides who may perform the underlying operation. A hidden button, a cached flag, or a client-supplied targeting key is not an access check.

Scope SDK credentials to the environment they can read, and enforce dashboard publishing permissions on the server. Explain that revoking a credential stops future fetches; it does not remotely erase configuration already cached by an offline process.

The portfolio moment

Two instances. One rollback. Different evidence.

Use this sequence as a demo script. These are proposed outcomes under the cache policy above, not measurements from a running service.

Instance B loses connectivity before revision 42 is published. The demo’s fallback is false.
CheckpointPublished revisionInstance AInstance B
Both connected41Evaluates 41Evaluates 41
Increase rollout42Adopts 42Uses cached 41 within age limit
Roll back43, with rules from 41Adopts 43Still on cached 41; adoption unknown centrally
B’s cache expires43Evaluates 43Returns default false, reason expired
B reconnects43Evaluates 43Validates and adopts 43

A matching value does not establish a matching version. B can coincidentally serve the rollback’s old rules while having no knowledge that a rollback happened. Its local debug view can show the cache; the central dashboard needs fresh evidence before claiming adoption.

The engineering bar in this product

Turn each capability into a repeatable check.

01 / Authentication & credential lifecycle

Build

Use established session authentication for the dashboard. Give each demo application instance a read-only, environment-scoped SDK credential. Keep it on the application server.

Demonstrate

Expire an admin session and rotate an SDK credential. Further fetches with the old credential fail; document how long a disconnected instance can still use previously fetched data.

02 / Authorization

Build

Separate viewers, editors, and publishers. Enforce workspace and environment permissions on every mutation, snapshot read, and audit-history request.

Demonstrate

An editor cannot publish to production, and a credential for staging cannot fetch production configuration. Feature visibility never substitutes for application authorization.

03 / Persistence & atomic publication

Build

Persist drafts, immutable environment snapshots, a current-version pointer, audit entries, and publish work. Commit the new version, pointer, audit entry, and notification work together.

Demonstrate

Restart after a publish acknowledgement. The same version is current, its actor and change are recorded, and unfinished notification work can still be found.

04 / Background jobs

Build

Use durable jobs for scheduled publications and update notifications. Include an expected current revision, an operation identity, and expiring worker claims.

Demonstrate

Stop a worker after publication commits. Recovery does not publish twice. A scheduled change whose base revision is stale is marked conflicted instead of overwriting an emergency rollback.

05 / Retries & recovery

Build

Fetch snapshots with deadlines and capped backoff with jitter. Keep evaluation independent of network requests. Make notification retries bounded; periodic refresh remains a catch-up path.

Demonstrate

Return 500 or 429 from the configuration endpoint. Refresh attempts spread out, evaluations return promptly, and recovery installs the current snapshot without a synchronized retry burst.

06 / Idempotency & concurrent edits

Build

Require a workspace-scoped operation key for publication and rollback. Atomically bind the key to the request parameters and result. Compare the expected current revision before publishing.

Demonstrate

Retry the same publication after its response is lost: it returns the original version. Reusing its key for different content conflicts. Two edits against the same base cannot silently overwrite each other.

07 / Migrations & compatibility

Build

Version stored rules and the snapshot schema. Backfill existing flags before enforcing new constraints, and keep a declared compatibility window for older SDKs.

Demonstrate

Upgrade a populated database and fetch with the previous SDK version. An unsupported snapshot is rejected as a whole; it never becomes an accidentally empty set of flags.

08 / Rate limiting & bounded work

Build

Limit mutation and snapshot requests per workspace and credential. Bound notification queues, SDK refresh concurrency, and telemetry buffers; coalesce version notices when only the latest snapshot matters.

Demonstrate

Reconnect many demo instances with a bounded test. Show admitted and rejected requests, stable worker concurrency, and continued service for another workspace.

09 / Observability & decision explanations

Build

Expose the evaluated value, reason, snapshot revision, cache freshness, and rule or rollout bucket in a debug view. Correlate publications with SDK adoption; aggregate or sample evaluation telemetry.

Demonstrate

Explain why two instances served different values for the same user. Trace the publication and snapshot versions without putting user identifiers in metric labels or logging every evaluation.

10 / Monitoring & rollback visibility

Build

Track publication failures, refresh errors, active SDK version lag, stale/default evaluation rates, and telemetry gaps. Define alerts and a runbook for a declared workload.

Demonstrate

Disconnect one instance during rollback. The dashboard distinguishes publication committed, adoption observed, and adoption unknown. Show the alert clearing after confirmed recovery.

11 / Deterministic rollout

Build

Specify a stable hash and encoding of environment, flag, rollout salt, and targeting key. Keep revision and percentage out of the hash. Compare a fixed bucket with the configured threshold.

Demonstrate

Run fixed test vectors and the same synthetic users through both SDK instances. Increasing the percentage retains the earlier cohort when identity, salt, and targeting rules stay unchanged.

Build → break → explain

Make the configuration unreliable on purpose.

Build a local harness that can delay or reject refreshes, reorder notifications, interrupt workers, and return invalid snapshots. Show which instance is affected and keep the fault state visible. Use synthetic users and a bounded client count.

Choose a failure drill

Availability & freshness

Configuration outage

Inject this fault
Block configuration fetches for one running instance, then start a second instance with no cached snapshot.
Predict before you run it
Do a warm instance and a cold instance have enough information to make the same decision?
Behavior to build
The warm instance may use its last validated snapshot within the declared age limit. Once that limit expires, it returns the caller’s default. The cold instance returns the default until initialization succeeds.

Evidence to collect

  • Evaluation latency stays independent of the failing network request.
  • The debug view distinguishes cached, expired, and not-ready results.
  • A failed refresh never resets the cache freshness timer.

Recover & explain

Restore connectivity and adopt a validated current snapshot. Record the recovery time and which requests used a fallback.

This planner describes drills to implement in your project. Selecting a drill does not send traffic or inject a fault.

What the chaos controls should expose

Let the reviewer select an instance, a fault profile, and a bounded duration or request count. Display the current published revision beside each instance’s local revision, cache age, decision reason, and last observed report.

Provide a stop action that restores ordinary transport, plus a reset that recreates synthetic fixtures with a clear run identity. Keep the incident’s evidence available. A deliberately broken evaluator belongs in the test harness so cohort drift is observable without changing the normal SDK contract.

A buildable sequence

Grow from a pure decision to an operated service.

  1. 01

    Make one decision explainable

    Build a pure boolean evaluator, fixed targeting fixtures, and a demo app with two visible feature paths. Return value, reason, and revision.

    Done when: The same snapshot and targeting key produce the same result. Off, explicit targeting, percentage boundaries, and missing context have documented behavior.

  2. 02

    Publish with an audit trail

    Add workspaces, environments, roles, drafts, immutable snapshots, atomic publication, and revision conflict checks.

    Done when: Unauthorized publications fail. Concurrent edits cannot silently overwrite each other, and a repeated publish request returns its original result.

  3. 03

    Give the SDK a lifecycle

    Create a small server-side TypeScript client with initialization, background refresh, atomic snapshot swaps, defaults, freshness tracking, and shutdown.

    Done when: Two instances stay responsive through outages, reject invalid or older snapshots, and explain cached versus default decisions.

  4. 04

    Schedule, roll back, and operate

    Add durable scheduled changes, revision-aware rollback, bounded retries, quotas, adoption telemetry, alerts, and a populated-data migration.

    Done when: A stale scheduled increase cannot overwrite a rollback. You can distinguish published state from observed SDK adoption and explain missing evidence.

  5. 05

    Package the failure story

    Expose fault profiles in a local demo harness, with synthetic users, a bounded client generator, clear reset steps, and a short incident report.

    Done when: A reviewer can reproduce a partitioned rollback and cohort-drift failure, inspect the evidence, and restore the system.

The portfolio handoff

Give the reviewer a decision they can investigate.

  1. Start with one user. Show the targeting key, stable bucket, value, and revision in both instances.
  2. Increase the rollout. Keep the existing cohort and explain who joins. Show the fixed fixture that verifies it.
  3. Partition one instance. Publish and roll back a change while its SDK cannot refresh.
  4. Inspect the disagreement. Trace the publication, installed revisions, age limit, and fallback. Distinguish local knowledge from central telemetry.
  5. Recover and qualify. Reconnect, verify adoption, and explain why this is neither instantaneous global rollback nor reversal of effects already performed.

Include the evaluator contract, test vectors, permission checks, a migration rehearsal, and a small incident report. Label propagation and latency measurements with the client count, workload, and environment. A debug display makes the reasoning inspectable; it does not replace verification.

Lessons behind this brief

Related lessons

The evaluator contract and the rollback drills build on these lessons.

Primary references

Use existing contracts to sharpen your own.

  • OpenFeature: flag evaluation API

    Defines caller defaults and detailed evaluation results. Use it to examine the shape of your client API; this brief does not claim OpenFeature compatibility.

  • OpenFeature: evaluation context

    Explains the targeting key and the context supplied to evaluation. Decide which identity your application owns and how it stays stable.

  • Unleash: stickiness

    A concrete implementation’s approach to consistent rollout assignments and grouping. Compare its choices with your fixture and hash contract.