← Architecture
Client and server A server shaped for one screen

Backend for frontend

One screen, one server that answers for it.

You have probably written one already: a +page.server.ts load or a route handler that calls a few APIs and hands the page what it needs. Ask an agent for a parcel-tracking page and it will write one too. Let’s watch what that server waits for, and what it sends.

TypeScriptGoOne tracking page, two routes, two recorded builds.

01 / The prompt

“Build me a parcel tracking page.”

A courier has three internal services. Shipments knows the parcel: where from, where to, which service. Scans knows every time a depot scanned it. Estimates predicts when it will arrive. You ask for the page a customer opens from the link in their email, and what comes back works: the parcel, its journey, a delivery window.

Somewhere in that build is a server that calls the three services and hands the page the result. That server is the subject of this lesson. The prompt never said what it should wait for, what it should do when one service is slow or down, or which of the services’ fields a customer should see.

The pattern has a name because teams kept meeting the same problem. Sam Newman describes it as having, instead of a general-purpose API, “one backend per user experience,” a name he credits to Phil Calçado at SoundCloud. On failure he asks: “if only the Inventory service was down, wouldn’t it be better to just degrade the functionality we pass back to the client?” (Pattern: Backends For Frontends, checked 23 September 2026). That question is the one the prompt left open.

02 / Name the shape

The page’s server answers for the page.

A backend for frontend is a server-side layer that belongs to one screen or one client. It calls the services behind it, and it owns the contract with that screen: which request the browser makes, what comes back, how long the screen waits, and what it shows when a part is missing. It does not own the facts. Shipments still decides where the parcel is going; the backend for frontend decides how the page says it.

The services decide what is true. The page’s server decides what the page waits for, what it shows, and what never leaves the server.

Who owns each part of the tracking page
WhatOwnerWhy
Where the parcel is going, its service levelShipmentsIt is the record. The route copies two fields and never computes them.
Each scan and its codeScansDepots write them. The route reads them and never edits them.
“Out for delivery” instead of OFDBackend for frontendWords are a screen decision. One map, with a fallback, in one place.
How long the page waits for each partBackend for frontendOnly the screen knows what it can live without.
What the page says when a part is missingBackend for frontendA named gap, not a blank page and not a guess.
Service tokens and internal idsBackend for frontendThey stay on the server. The reply carries only what the screen shows.
Rendering, and asking againBrowserIt makes one request, shows each state the reply can be in, and offers a reload.

Words to put in a prompt or a review

Backend for frontend (BFF)
A server layer owned by one screen or client, shaped for it.
Fan-out
One incoming request that becomes several calls to other services.
Deadline
How long one call may take before it is cancelled and treated as missing.
Required part
Without it there is no page, so its failure is the page’s failure.
Partial response
A reply that says what it has, and names what it does not have and why.
Pass-through
A route that forwards what the services said, in their shapes.
View model
The data a screen renders, in the screen’s words and nothing more.
How many backends for frontends?One per experience, and where it lives

Newman’s rule of thumb, on the same page, is “one experience, one BFF”: if the iOS and Android apps are very similar, one BFF can serve both, and a website with a different job gets its own. In this lesson the app and the website share one route with two views, chosen by a query parameter; when the two screens start to pull the route in different directions, split it.

It does not have to be a separate service. A SvelteKit +page.server.ts load or a Next.js server component that calls three APIs is a backend for frontend for one page. A GraphQL layer in front of the services is another way to let a screen ask for its shape; it moves the shaping into the query and leaves the deadline and missing-part questions where they were.

03 / Follow one page load

Watch the same page wait for two different things.

First the pass-through route: it asks all three services and waits for every answer. Then the backend for frontend: each part has a deadline, and a missing part is named instead of failing the page. The estimate service is the one that misbehaves. Open Try it and choose what each service does today.

Backend for frontend

What does the page wait for?

Pass-through route · 0 ms

Shipmentswaiting

Scanswaiting

Estimatewaiting

What the browser receives

Waiting for the route…

01/ 05
Forward everything

The pass-through route asks all three and waits for every answer.

Three requests leave at once. The page waits for the route, and the route waits for all three.

Reduced motion: choose a scene to see its completed state.

Read this scene

Three requests leave at once. The page waits for the route, and the route waits for all three.

Pass-through route at 0 ms. Shipments: waiting; Scans: waiting; Estimate: waiting. The browser is still waiting.

Watch restarts the story when you come back. Step through keeps your step. Try it runs the two routes fresh each time you load the page.

04 / Read the shape

A deadline per part, and a reply per screen.

Basic form is one call under one deadline, with an outcome that has a name. In the wild is the route: which part is required, which parts may be missing, and the reply shaped for the screen that asked, beside the pass-through route it replaces. At the call site the browser makes one request and the server’s handler answers it.

Notice what trackingPage awaits first: only Shipments, the part without which there is no page. The other two calls started at the same moment and are still running under their own deadlines.

One dependency under one deadline, with an outcome that has a name: a value, or timed-out, failed, or not-found. When the deadline passes the call is cancelled, not left running.

TypeScriptReading
tracking.ts
export type Outcome<T> =
	{ ok: true; value: T } | { ok: false; reason: 'timed-out' | 'failed' | 'not-found' };

// One dependency, one deadline, and an outcome with a name. When the deadline
// passes, the call is cancelled, not left running.
export async function withDeadline<T>(
	clock: Clock,
	limit: number,
	parent: AbortSignal,
	call: (signal: AbortSignal) => Promise<T>
): Promise<Outcome<T>> {
	const request = new AbortController();
	const stop = () => request.abort(parent.reason);
	parent.addEventListener('abort', stop, { once: true });
	const timer = new AbortController();
	try {
		return await Promise.race([
			call(request.signal).then((value): Outcome<T> => ({ ok: true, value })),
			clock.sleep(limit, timer.signal).then((): Outcome<T> => {
				request.abort(new Error('deadline'));
				return { ok: false, reason: 'timed-out' };
			})
		]);
	} catch (error) {
		return {
			ok: false,
			reason: error instanceof ServiceError && error.status === 404 ? 'not-found' : 'failed'
		};
	} finally {
		timer.abort();
		parent.removeEventListener('abort', stop);
	}
}
GoAlongside
main.go
// Outcome is a value, or the name of why there is none.
type Outcome[T any] struct {
	Value  T
	Reason string // "", "timed-out", "failed", or "not-found"
}

// WithDeadline starts one dependency under its own deadline. When the deadline
// passes, the call's context is cancelled, not left running.
func WithDeadline[T any](ctx context.Context, limit time.Duration, call func(context.Context) (T, error)) <-chan Outcome[T] {
	done := make(chan Outcome[T], 1)
	go func() {
		ctx, cancel := context.WithTimeout(ctx, limit)
		defer cancel()
		value, err := call(ctx)
		var failure *ServiceError
		switch {
		case err == nil:
			done <- Outcome[T]{Value: value}
		case errors.Is(err, context.DeadlineExceeded):
			done <- Outcome[T]{Reason: "timed-out"}
		case errors.As(err, &failure) && failure.Status == 404:
			done <- Outcome[T]{Reason: "not-found"}
		default:
			done <- Outcome[T]{Reason: "failed"}
		}
	}()
	return done
}
The behavior these examples promiseChecked by 12 shared scenarios, each run through both routes
  • All three calls start at once. Shipments and Scans have 800 ms each, the estimate 300 ms. A call past its deadline is cancelled and counts as timed-out.
  • No parcel from Shipments, for any reason, means no page: a 404 when Shipments says there is no such parcel, otherwise 502 with part: "shipment". The other calls are cancelled at that moment.
  • A missing Scans or estimate gives a 200 with status: "partial" and a missing entry naming the part and the reason. The page answers when the last part has answered or run out of time.
  • The reply has the parcel’s id, route, and service; the newest scan in words; the full history only for the website; the delivery window. No account, staff id, recipient, confidence, or model name.
  • A scan code the route does not know reads “Update at <depot> depot”.
  • The pass-through route answers with the three payloads as sent, or 502 as soon as any call fails, leaving the others running.

Every expectation in the shared cases, down to the millisecond each route answers and the state each call was in, was produced by a separate model written from these rules and kept beside the examples, not copied from either implementation. The tests run on a clock they control, so the times are exact.

Reading the TypeScriptAbortController, Promise.race, and a clock

withDeadline races the call against a timer and aborts the call’s AbortController when the timer wins, so a well-behaved client such as fetch closes the connection. A second controller for the whole page aborts every call at once when Shipments fails. The clock is passed in so the tests can run the same code on virtual time; in production it is setTimeout.

Reading the Gocontext.WithTimeout, and synctest

Each part runs in a goroutine under context.WithTimeout, and returns its outcome on a buffered channel so it never blocks after the route has stopped listening. defer cancel() on the page’s context is the whole cancellation story. The tests run inside testing/synctest, where timers use a fake clock, so 1.5 seconds takes no time and elapsed times are exact.

Run it yourselfNo dependencies

Copy the complete TypeScript file and run node --experimental-strip-types tracking.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print:

go.mod
module heyrian.dev/lessons/backend-for-frontend

go 1.25
bff app · 200 complete · Out for delivery, Sep 23 07:48 · arriving 10:00–12:00 · 203 bytes
bff app, slow estimate · 200 partial · Out for delivery, Sep 23 07:48 · no estimate (timed-out) · 226 bytes
pass-through, slow estimate · 200 pass-through · 3 raw payloads · 736 bytes
pass-through, estimate down · 502 upstream failed
bff web, new scan code · 200 complete · Update at Leeds depot, Sep 23 08:30 · arriving 10:00–12:00 · 7 scans listed · 483 bytes

05 / Review the agent’s diff

“The estimate always loads now.”

Customers seeing “we can’t estimate a delivery time” is a real complaint, and the fix is one number. Read what that number controls before you decide.

The agent’s pull request

“Customers kept seeing ‘Estimate unavailable’. Raised the estimate’s deadline so it always loads. All 31 tests pass.”

// server/tracking.ts
			export const deadlines = {
			  shipment: 800,
			  scans: 800,
			(removed)  eta: 300
			(added)  eta: 5000 // estimates kept timing out
			};
			
You are reviewing this change. What do you do?

06 / How it fails

Three services give the page three ways to be late.

Every row is one of the shared scenarios, run through both routes. The times are the route’s own, from the tests’ virtual clock.

Failure modes of one tracking page load
What goes wrongWhat the customer seesWhat the backend for frontend doesPass-through
Slow optional part: the estimate takes 1.5 sThe parcel at 300 ms, and “we can’t estimate a delivery time right now”.Cancels the estimate call at its deadline and names the gap.Answers at 1,500 ms, with everything.
Slow required part: Shipments takes 1 sAn error at 800 ms that says tracking is unavailable.Answers 502 at the deadline instead of waiting on.Answers at 1,000 ms, with everything.
Unreachable optional part: the estimate is downThe parcel, with the estimate marked as failed.Answers at 60 ms, when Scans is in.502 at 50 ms. No parcel. Scans is left running.
Unreachable required part: Shipments is downAn error page, at once.502 at 30 ms, and cancels the other two calls.502 at 30 ms, and leaves the other two running.
Half there: estimate slow and Scans downThe parcel, with two named gaps.200 partial at 300 ms, missing scans (failed) and eta (timed-out).502 at 30 ms.
Wrong: a scan code the route has never seen“Update at Leeds depot”.Falls back to where the parcel is, and keeps the rest of the page.Forwards HLD; each screen decides what to print.
Wrong: no parcel with that number“No parcel with that number”.404 at 25 ms, and cancels the other calls.502, as if a service broke.
Duplicated: the customer reloadsThe same page again.Nothing to guard: every call is a read. Not modeled in the cases.The same.

Two ideas carry the table: a deadline is only half a timeout unless the call is cancelled too, and a partial answer is honest only when it names what is missing. Both have their own lessons: Timeouts, deadlines, and races, Cancellation propagation, and Aggregate and partial failure.

07 / Is it worth it?

You pay for a layer. Here is what it buys.

The pass-through route is shorter and has nothing to configure. Hold both up against the kinds of change every app gets.

The same four changes, made to each route
ChangePass-throughBackend for frontend
A second client: the website wants the full historyNothing to change on the server. Each client reshapes 736 bytes itself and carries its own copy of the code-to-words map.A second view in one place: 438 bytes for the website, 203 for the app.
Replace a dependency: Estimates ships a new response shapeEvery client reads the new shape, including app versions already on phones.One change in the route. The screens’ contract does not move.
Change a rule: hide estimates the model is unsure ofEach client has to learn it, and old app versions never will.One condition in the route, deployed once.
A second team takes over the appThey depend on three services’ shapes and three teams’ release notes.They own their route. The risk moves: keep domain rules, such as whether a parcel is late, in the services and out of the route.

Before you add the layer, write down what you will measure and the result you would accept, so it is judged by something other than how the diagram looks:

  • Time until the parcel is on screen, at the 95th percentile, measured in the browser, before and after. This is what the deadlines are for.
  • Share of pages answered partial, by part and reason. If the estimate misses 300 ms on a large share of loads, fix the estimate service or stream it; do not raise the deadline in the dark.
  • Bytes per page on the app, and calls still running after the page answered. The second should be zero.

This page did not run the tracking page for real customers, so it has no production numbers. The checker’s timings in section 08 are one sample on one machine, not a benchmark.

08 / Ask for it

Two prompts, two builds, one browser.

We sent two agents the same request for this page at the same time, both running Claude Sonnet, with the three services running for them to call. One prompt described the page. The other added an Architecture block: one data request shaped for the page, no field the page does not show, Shipments required, a deadline per call that cancels it, and named gaps. Then a script ran each build behind services it controlled and opened the page in Chromium, once per question.

What the checker found, run 2026-09-23
QuestionPlain promptArchitecture prompt
Data requests the browser makesNone: the server sends a finished page/api/track/PX-4471 200 669 bytes
Internal ids a customer can read on the pageacct_77, u-2291, u-3307None
How the three calls startOne after anotherShipments first, then the other two together
The estimate takes 1.5 s: parcel shown after1,659 ms382 ms
The estimate never answers: parcel shown after4,164 ms377 ms
Scans take 1.2 s: parcel shown after1,407 ms878 ms
The estimate is down: what the page saysEstimated arrival is unavailable right now.We couldn't load the delivery estimate right now (the service failed).
Scans send a code neither prompt listed, HLDHLDStatus update (HLD)
Shipments is downTracking is temporarily unavailableCould not look up parcel "PX-4471": the shipments service failed.

The plain prompt did not produce the browser calling three services. It produced a server that called them, held the tokens, and sent the browser one finished page: a backend for frontend in all but name. What it never decided was what to wait for. It called the services one after another, each with a four-second timeout, so the page waited for the estimate every time the estimate was slow. And it showed every field it was given, which put the sender’s account id and the depot staff ids on a customer’s phone.

The architecture build did what its block asked, and the checker could see each part of it: one data request, no internal field in the reply, the estimate call hung up at 300 ms, and a page that says which part is missing. It also did one thing the block never forbade: it waited for Shipments before starting the other two calls, so every load pays Shipments’ time first.

server.ts · plain prompt
let scans: Scan[] = [];
let scansUnavailable = false;
try {
  scans = await getScans(parcelId);
} catch (err) {
  console.error("scans lookup failed:", err);
  scansUnavailable = true;
}

let eta: Eta | null = null;
let etaUnavailable = false;
try {
  eta = await getEta(parcelId);
} catch (err) {
  console.error("eta lookup failed:", err);
  etaUnavailable = true;
}
server.ts · architecture prompt
async function handleTrackApi(id: string): Promise<TrackApiResult> {
  const shipmentsResult = await callService(
    `${SHIPMENTS_URL}/shipments/${encodeURIComponent(id)}`,
    SHIPMENTS_AUTH,
    SHIPMENTS_TIMEOUT_MS,
  );

  if (!shipmentsResult.ok) {
    if (shipmentsResult.status === 404) {
      return { status: 404, body: { error: `No parcel found for "${id}".` } };
    }
    return {
      status: 502,
      body: {
        error: `Could not look up parcel "${id}": the shipments service ${shipmentsResult.reason}.`,
      },
    };
  }

  const [scansResult, etaResult] = await Promise.all([
    callService(`${SCANS_URL}/scans/${encodeURIComponent(id)}`, SCANS_AUTH, SCANS_TIMEOUT_MS),
    callService(`${ETA_URL}/eta/${encodeURIComponent(id)}`, ETA_AUTH, ETA_TIMEOUT_MS),
  ]);

Two lines were missing from both prompts. One is timing: start every call at once; nothing waits for another part unless it needs that part’s answer. The other is hidden in a product sentence. “Sees the parcel, where it has been, and when it should arrive” reads like a feature, and the plain build took it to mean every field of the parcel. The line that closes it names the fields: the page shows these fields and no others.

Neither build hid a code it did not know. The plain build’s status badge showed HLD; the architecture build’s fallback kept it in brackets. The prompt asked for a “sensible fallback” and did not say what a customer should read.

How the runs were made and checkedOne run each, recorded as written
  • Both agents received the prompts word for word, in fresh contexts, in the same message. Neither was told about the other, this lesson, or the checker. The only differences were the Architecture block and the output folder.
  • The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s examples. For every question the checker restores a build into a temporary folder, starts its own fake services with that question’s behavior, and opens the page at phone size in Chromium, recording every response the browser receives.
  • The checker’s first run reported the plain build’s estimate message as “Estimated arrival”, which is the section’s heading. It now reads the line after the heading too. Both runs are kept.
  • Both agents wrote something outside their folders against the prompt: the plain agent saved two test pages to a system temporary folder, and the architecture agent tried to write a log at the filesystem root, which failed. Neither touched the other’s folder.
  • This is one sample of each prompt, not a measurement of a model. The timings are one machine’s, a few milliseconds either way between checker runs.

09 / Hold it there

Keep the tokens behind the door, and the deadlines under test.

The next change can quietly undo this: a component that fetches Scans directly because it was quicker, a deadline raised to make a complaint go away. Three kinds of check keep the shape.

  1. The framework’s own door

    Service tokens belong in server-only configuration. SvelteKit refuses to let code that reaches the browser import $env/static/private, $lib/server, or *.server.ts files, and names the import chain when it does (server-only modules). In Next.js, import 'server-only' makes importing a module into a Client Component a build error, and environment variables without the NEXT_PUBLIC_ prefix are empty strings in the browser (Server and Client Components).

  2. An import rule an agent cannot argue with

    Only the page’s server route, and the server folder it lives beside, may import the service clients. A component that wants Scans asks the route. Enforcement layer runs rules like this against real code, and Architecture as rules writes them from one declaration. This rule was not run against this lesson’s files, which have no app around them.

    .dependency-cruiser.cjs
    // .dependency-cruiser.cjs
    module.exports = {
    	forbidden: [
    		{
    			name: 'only-the-page-route-calls-services',
    			comment: 'Service clients hold tokens. Components and browser code ask the route instead.',
    			severity: 'error',
    			from: { path: '^src/', pathNot: '\\+page\\.server\\.ts$|^src/lib/server/' },
    			to: { path: '^src/lib/server/services/' }
    		}
    	]
    };
  3. A check on what actually happens

    Rules see imports, not timing and not payloads. So check the behavior: open the page with a slow service behind it and time it, and search everything the browser received for tokens and internal ids. The checker in section 08 does both, and the lesson’s own tests assert that the estimate call is cancelled at 300 ms and that no internal field is in the reply.

    check-runs.mjs
    /** Open a page in Chromium and record everything the browser receives. */
    async function visit(browser, path, { waitFor = 'Leeds', limit = 12_000 } = {}) {
    	const page = await browser.newPage({ viewport: { width: 390, height: 844 } });
    	const received = [];
    	page.on('response', async (response) => {
    		const request = response.request();
    		let body = '';
    		try {
    			body = await response.text();
    		} catch {
    			/* redirects and aborted bodies have none */
    		}
    		received.push({ url: response.url(), type: request.resourceType(), status: response.status(), bytes: Buffer.byteLength(body), body });
    	});
    	const started = Date.now();
    	const document = await page.goto(`${BASE}${path}`, { waitUntil: 'commit', timeout: limit });
    	let shownAfter = null;
Your load function is already oneA page’s server load that calls several APIs is a backend for frontend. The deadline is the part you own.

Where it already is in your components

A SvelteKit +page.server.ts load that calls three APIs and returns what the page renders is a backend for frontend for one page. So is a Next.js server component that awaits three fetches before it returns markup. They run on the server, so they can hold tokens; they decide the shape the component receives; and whatever they await is what the page waits for.

Most of them are written like the pass-through route: one Promise.all, every field passed down. That is fine until one of the three is slow.

When you have to own it

The day the estimate service slows down, you decide what the page does without it. Keep the required parts under a deadline and render them, and send the optional part separately. SvelteKit streams a promise a server load returns without awaiting, “useful if you have slow, non-essential data, since you can start rendering the page before all the data is available” (Streaming with promises); streaming needs JavaScript in the browser. In React, the same move is a <Suspense> boundary around the component that reads the estimate.

The tracking page as most of us write it: the server calls the three services, and the component renders the parcel, the latest scan, and the estimate, with a message for a missing part.

ReactAlready in your code
app/track/[id]/page.tsx
// app/track/[id]/page.tsx: a Next.js server component. It runs on the server,
// so it can hold service tokens and call three internal APIs. The browser gets
// only the HTML this returns.
import { notFound } from 'next/navigation';
import { shipments, scans, estimates } from '@/server/services';

export default async function TrackingPage({ params }: { params: Promise<{ id: string }> }) {
	const { id } = await params;
	const [parcel, history, estimate] = await Promise.allSettled([
		shipments.get(id),
		scans.list(id),
		estimates.get(id)
	]);
	if (parcel.status === 'rejected') notFound();

	const latest = history.status === 'fulfilled' ? history.value.at(-1) : undefined;
	return (
		<main>
			<h1>Parcel {parcel.value.id}</h1>
			<p>
				{parcel.value.from} → {parcel.value.to} · {parcel.value.service}
			</p>
			<p>
				{latest ? `${latest.text}, ${latest.time}` : 'Tracking history is unavailable right now.'}
			</p>
			<p>
				{estimate.status === 'fulfilled'
					? `Arriving ${estimate.value.window}`
					: 'We can’t estimate a delivery time right now.'}
			</p>
		</main>
	);
}

10 / Make the call

Add the layer when a screen has something to decide.

Let the browser call a service directly when there is one service, it was built for browsers, its answer is what the screen shows, and there is no token to hide: a public read API behind a CDN, say. Then a layer in between only adds a hop. That is plain client–server.

Add a backend for frontend when the screen needs several services, when one of them can be slow and the screen can live without it, when fields must not reach the browser, or when a second client wants a different shape. Reopen the decision when the route starts deciding things that are true for every client, such as whether a parcel is late: that rule belongs in a service.

Take it with you

Explain it without saying “backend for frontend”: “The page asks its own server once. That server calls the others at the same time, gives each a time limit, and tells the page what it got, in the page’s words, and what it could not get.” Then open the last page load you wrote that calls more than one API, and find what it awaits.

Paste into your next prompt, and fill in the blanks

The browser makes one request for <this screen>'s data, to one endpoint
shaped for it. It never calls <the services> and never receives their tokens.
The reply carries only the fields the screen shows: <list them>.
<The required part> is required; without it, answer with a clear error.
<The optional parts> are optional.
Start every call at once. Give each a deadline: <n> ms for <part>.
When a deadline passes, cancel that call and answer without it.
The reply names each missing part and why: timed out or failed.
The services own the facts; this route only chooses, translates, and shapes them.
Connections to follow nextRelated lessons

Take the route into your editor. When the newest scan says the parcel was delivered, skip the estimate call entirely, and decide what the page shows in its place.

Back to architecture →