← Architecture
Failure and evidence Finding one request

Observability across a request

One id, every hop.

Every request you make crosses more than one program: your server, the one it calls, the one that calls. When something goes wrong, each of them wrote down its own part. Let’s take one support ticket, a busy minute of logs, and find out whether those parts can be put back together.

TypeScriptGoOne unlock, four services, two recorded builds.

01 / The prompt

“Build the unlock flow, and log each request.”

A city bike-share: a rider taps Unlock, a gateway calls the rides service, rides puts a $1 hold on the card through payments and asks the dock controller to release the bike. You ask an agent for it, with logging. What comes back works, and every service writes neat JSON lines.

Then a ticket arrives: “r-88, about 08:14, bike B-12. The app said something went wrong, and there’s a $1 charge on my card.” The gateway log has one line for r-88: a 502 after two seconds. Rides, payments, and the dock each have lines from 08:14, and eleven other riders were unlocking bikes in the same minute. Which lines are r-88’s?

The prompt said to log each request. It did not say that a request is one thing across four programs, so each program logged its own. That is the question it never answered: how does one request keep one identity across every hop?

02 / Name the shape

One id, every hop.

Observability across a request means one request can be followed through every service it touches, after the fact, from whatever the user can tell you. The mechanism is context propagation: the first service gives the request an id, and every service passes it on, on every call, and writes it on every line.

A request’s identity is created once, at the edge, and inherited everywhere else. No service invents its own.

Here is who knows what about r-88’s unlock without it.

What each service knows about one unlock, without a shared id
ServiceKnowsCannot tell you
GatewayThe rider, the bike, the status, how long it tookWhat happened inside.
RidesThe rider and bike, the hold, each dock attemptWhat the dock did after it stopped waiting.
PaymentsAn amount and a cardWhich unlock the hold was for.
Dock controllerA bike and a stationWho asked, or whether anyone was still listening.

Words to put in a prompt or a review

Trace
Everything one request did, across every service, under one id.
Span
One service’s part of it, with its own id and its caller’s id as its parent.
Context propagation
Passing the trace id and the current span id on every call.
traceparent
The W3C header that carries them: 00-traceid-spanid-flags.
Correlation
Joining lines by a shared id, not by time or by guessing from fields.
Support reference
A short piece of the trace id shown to the user when something fails.
Logs, metrics, and tracesWhich one answers which question

Metrics say how often and how slow, for everyone: a rising 502 rate at 08:14. Logs say what one service did. Traces say what one request did everywhere. The trace id is what lets a log line and a trace point at the same request; tools like OpenTelemetry record spans and attach the trace id to logs for that reason. You can start with logs alone, as this lesson does: the propagated id is the part that matters, and a tracing backend can come later.

03 / Follow one unlock

Twelve riders, one slow dock, and a ticket.

First the minute as it was logged with service-local ids, and support searching it by time. Then the same minute with a trace id passed on every call, and support searching by that id. Open Try it to search for any rider yourself.

Failure and evidence

Which of these lines are r-88’s?

Service-local request ids · 08:14:00.000

gateway.log

  1. No lines yet

rides.log

  1. No lines yet

payments.log

  1. No lines yet

dock.log

  1. No lines yet

Replaying the logs as they were written

01/ 04
Twelve riders, one minute

Twelve riders tap Unlock, 150 ms apart.

Each service writes its own log line with its own request id. Nothing ties one service’s id to another’s.

Reduced motion: choose a scene to see its completed state.

Read this scene

Each service writes its own log line with its own request id. Nothing ties one service’s id to another’s.

Service-local ids, 0 ms after 08:14:00. gateway: nothing.rides: nothing.payments: nothing.dock: nothing.

Watch restarts the story when you come back. Step through keeps your step. Try it lets you search the same minute as support.

04 / Read the shape

A header, a span, and a call that passes it on.

Basic form is the header format, and what to do with a bad one. In the wild is the span each service opens: continue the caller’s trace, or start one. At the call site the context crosses the wire, and every outgoing call carries it.

The W3C traceparent header: version, a 32-hex trace id, the caller’s 16-hex span id, and flags. Anything malformed, uppercase, or all zeros is no context, and the service starts a new trace.

TypeScriptReading
trace.ts
/** One hop's view of a request: which trace it belongs to and which span called it. */
export type TraceContext = { traceId: string; spanId: string; sampled: boolean };

const TRACEPARENT = /^([0-9a-f]{2})-([0-9a-f]{32})-([0-9a-f]{16})-([0-9a-f]{2})$/;

/** Read a `traceparent` header. Anything malformed, or all zeros, is no context at all. */
export function parseTraceparent(header: string | null | undefined): TraceContext | null {
	const match = header ? TRACEPARENT.exec(header.trim()) : null;
	if (!match) return null;
	const [, version, traceId, spanId, flags] = match;
	if (version === 'ff' || /^0+$/.test(traceId) || /^0+$/.test(spanId)) return null;
	return { traceId, spanId, sampled: (parseInt(flags, 16) & 1) === 1 };
}

export function formatTraceparent(context: TraceContext): string {
	return `00-${context.traceId}-${context.spanId}-${context.sampled ? '01' : '00'}`;
}
GoAlongside
main.go
// TraceContext is one hop's view of a request: which trace it belongs to and which span called it.
type TraceContext struct {
	TraceID string `json:"traceId"`
	SpanID  string `json:"spanId"`
	Sampled bool   `json:"sampled"`
}

var traceparent = regexp.MustCompile(`^([0-9a-f]{2})-([0-9a-f]{32})-([0-9a-f]{16})-([0-9a-f]{2})$`)

// ParseTraceparent reads a traceparent header. Anything malformed, or all zeros, is no context.
func ParseTraceparent(header string) *TraceContext {
	m := traceparent.FindStringSubmatch(strings.TrimSpace(header))
	if m == nil || m[1] == "ff" || strings.Trim(m[2], "0") == "" || strings.Trim(m[3], "0") == "" {
		return nil
	}
	flags, _ := strconv.ParseUint(m[4], 16, 8)
	return &TraceContext{TraceID: m[2], SpanID: m[3], Sampled: flags&1 == 1}
}

func FormatTraceparent(c TraceContext) string {
	flags := "00"
	if c.Sampled {
		flags = "01"
	}
	return "00-" + c.TraceID + "-" + c.SpanID + "-" + flags
}
The behavior these examples promiseChecked by 38 shared cases in TypeScript and Go
  • A traceparent is two lowercase hex digits, 32, 16, and 2, joined by dashes, with surrounding spaces ignored. Version ff, an all-zero trace id, or an all-zero span id makes it invalid. The sampled flag is the lowest bit.
  • A span continues a valid incoming trace, with the caller’s span as its parent, or starts a new trace with no parent. Its outgoing header names itself as the parent.
  • The simulated minute: twelve riders 150 ms apart; gateway to rides 5 ms; a hold after 80 ms; station 3’s dock answers in 120 ms and station 7’s in 3000; rides waits 1000 ms per attempt, twice, then answers 502 and keeps the hold; the dock logs its unlocks when they finish.
  • Support starts from the gateway line for the rider. With a trace id it takes every line in that trace; without one, every line written while the request was open.

Every expectation in the shared cases was produced by a separate model written from these rules, kept beside the examples in examples/model/, not copied from either implementation: twelve headers, both minutes line for line, and what support finds for each of the twelve riders in each shape.

Reading the TypeScriptA regex, a class, and Headers

One regular expression checks the whole header, so a header is either a context or null, never half of one. Span keeps the service and log sink in private fields; outgoing() is the only way to get a header out of it. new Headers(init.headers) in tracedFetch keeps whatever headers the caller set and adds one.

Reading the GoA pointer for no context, and a handler

ParseTraceparent returns *TraceContext, nil for no context. strings.Trim(id, "0") == "" is the all-zero check. The rides handler reads the header from r.Header and continues the trace before doing anything else.

Run it yourselfNo dependencies

Copy the complete TypeScript file and run node --experimental-strip-types trace.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print:

go.mod
module heyrian.dev/lessons/observability-across-a-request

go 1.23
local: r-88 got 502 after 2090 ms; found gateway 5, rides 11, payments 4, dock 4; the bike's unlock missed
traced: r-88 got 502 after 2090 ms; found gateway 1, rides 4, payments 1, dock 2; the bike's unlock found

The minute is simulated, and the ids count up so both languages print the same thing. Real services use random ids.

05 / Review the agent’s diff

“Adds request ids to the rides service.”

Structured logs with ids are a real improvement over printed strings. Read where this id comes from, and where it goes.

The agent’s pull request

“Adds structured logging with request ids to the rides service, so support can trace failed unlocks. All 24 tests pass.”

// rides/server.ts
			export async function handleUnlock(request: Request) {
			(added)  const requestId = crypto.randomUUID();
			(added)  const log = (msg: string, fields = {}) =>
			(added)    append("logs/rides.log", { requestId, msg, ...fields });
			  const { rider, bike, station } = await request.json();
			  const hold = await post(`${PAYMENTS_URL}/holds`, { rider, amountCents: 100 });
			(added)  log("hold placed", { rider, holdId: hold.holdId });
			  const released = await unlockWithRetry(bike, station);
			(added)  log(released ? "unlocked" : "unlock failed", { rider, bike });
			
You are reviewing this change. What do you do?

06 / How it fails

A trace breaks at the first hop that forgets it.

Each way the unlock can go wrong, what the rider sees, and what the traced build lets support see. The first row comes from the shared cases; the rest are authored from the same example.

Failure modes of one unlock, and what a trace shows
What goes wrongWhat the rider seesWhat support finds with a traceWithout one
Slow dock, unlock after the rider gave upAn error, a hold, and a bike that is free.8 lines, including the dock unlocking B-12 a second after the 502.24 lines, most of them other riders’; the late unlock missed
Payments downAn error, no hold.The rides span, a failed payments call, and no dock call.Payments’ own error, with no rider on it
A hop drops the headerNothing different.A trace that stops at that hop; the next service starts a new trace.No difference to notice
Two taps, two requestsOne unlock, maybe two holds.Two traces, each with its own hold.Two holds and no way to tell which tap made which
A malformed or all-zero header arrivesNothing.A new trace from that hop; the bad header is not trusted.Not applicable
A trace id in the user-facing errorA reference to read out.The whole request from that reference.A time and a guess

The slow dock is also a failure-mode question: an unlock whose outcome is unknown should not be reported as failed, and the hold should not outlive it. Tracing does not fix that; it is how you find out it happened. Keeping one slow dependency from slowing everything else is Containing failure.

07 / Is it worth it?

You pay one header and one field per line. Here is what they buy.

Service-local logs are simpler to write and are enough while one service does all the work. Hold both shapes up against the changes a system like this gets.

The same four changes, made to each shape
ChangeService-local idsPropagated context
A second client: the kiosk at each stationAnother source of lines to line up by time.The kiosk starts or sends a traceparent; its requests are traces like any other.
Replace the dock vendorNew log formats to learn to join.The new vendor receives the header. Whether it logs it is a question for the contract.
Change a rule: three dock attempts, not twoOne more line to find by time.One more span under the same trace.
A second team owns paymentsTheir logs never meet yours.A shared id is the only shared language the two teams’ logs need.

Before you add propagation, decide what you would measure and what you would accept:

  • Time to explain a ticket, from ticket to cause, for the next ten tickets. Take it before the change, and after.
  • Share of log lines with a trace id, per service. Anything under all of them is a hop that breaks the chain.
  • Traces with one span where you expected several: a hop that dropped the header, found by counting instead of by a bad day.

This page did not run a bike-share, so it has no numbers of that kind to give you. Which requests count as failures in the first place is Defining success.

08 / Ask for it

Two prompts, twelve riders, one ticket to follow.

We sent two agents the same request for this unlock flow at the same time, both running Claude Sonnet. Both prompts asked for JSON log lines in each service. One added a paragraph: support must take a ticket that names a rider and a time, find that one request, and find every line it left and every call it made, with one search, even when many riders unlock in the same second. It named no header and no standard. Then a script ran each build against fake payments and dock services, unlocked twelve bikes 150 ms apart with one slow dock, and followed whatever id r-88’s gateway line carried.

What the checker found, run 2026-09-23
What the checker looked atPlain promptArchitecture prompt
r-88’s unlock, in a minute of twelve502 after 2.1 s502 after 2.1 s
Ids on r-88’s gateway log lineNone1
Rides log lines linked by those ids0 of 37 (3 mention r-88)8 of 75 (8 mention r-88)
Calls to payments that carry them0 of 121 of 12
Calls to the dock that carry them0 of 13 (2 were for B-12)2 of 13 (2 were for B-12)
Header that carried themNonex-request-id

The plain build logged carefully and joined nothing. Its rides lines carry the rider, bike, station, and hold id, so a patient person could match them to the gateway by rider. But its gateway line for r-88 carries no id, and neither the payments call nor the two dock calls for B-12 carry anything that says which request they belong to.

The architecture build did what this lesson teaches, under another name. It made a UUID at the gateway, wrote it on every line with the rider, and sent it as X-Request-Id on every call, including both dock attempts. One search for that id found exactly r-88’s lines and calls.

server.ts · plain prompt
log(RIDES_LOG, { event: "unlock_requested", rider, bike, station });

const hold = await placeHold(rider);
if (!hold) {
  log(RIDES_LOG, { event: "hold_failed", rider, bike, station });
  sendJson(res, 502, { unlocked: false, error: "payment hold failed" });
  return;
}
log(RIDES_LOG, { event: "hold_placed", rider, bike, station, holdId: hold.holdId });

const released = await releaseBike(bike, station);
rides.ts · architecture prompt
const { status, body } = await postJson(
  `${paymentsUrl}/holds`,
  { rider: ctx.rider, amountCents: 100 },
  { "X-Request-Id": ctx.requestId },
);

Two details are worth the prompt line. The header was home-made, so a vendor, a proxy, or a tracing library would not recognize it; the W3C traceparent is the one they already pass on. And the id starts at the gateway, so the rider’s screen has nothing to read out to support. The line the runs point to: start a W3C trace context in the client, continue it in every service, and show the user a short reference when a request fails.

How the runs were made and checkedOne run each, recorded as written
  • Both agents received the prompts word for word, in fresh contexts, in the same message. Neither was told about the other, this lesson, or the checker. Besides the Architecture paragraph and the output folder, the prompts differed in one line: the ports each could use to test with, so their stand-ins could not collide.
  • The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s examples. The checker restores each build, starts it with fake payments and dock services that record every call, unlocks twelve bikes 150 ms apart with one slow dock, and follows any id on r-88’s gateway line.
  • This is one sample of each prompt, not a measurement of a model. The transcript audit and every number are in the run notes beside the examples.

09 / Hold it there

Make a missing trace id a failing test.

Propagation breaks quietly: a new call written with a bare fetch, a new service that logs with console.log. Three checks catch it.

  1. The standard, not a home-made header

    Use the W3C Trace Context format. Its traceparent is version-trace-id-parent-id-trace-flags, an all-zero trace id “is considered an invalid value”, and when the header is invalid “the vendor creates a new traceparent header” (W3C Trace Context). Tracing libraries, proxies, and vendors already speak it, so a hop you do not own may pass it on for you.

  2. A test that follows one request

    Send one request through the hops you own, with the calls to payments and the dock pointed at a local server that records headers, and assert every log line and every outgoing call carries the same trace id. This lesson’s suite does that for the gateway-to-rides hop.

    trace.spec.ts
    it('sends the current span as the parent, and the callee continues the trace', async () => {
    	const sink: LogLine[] = [];
    	let n = 0;
    	const ids: Ids = {
    		trace: () => 'cd'.repeat(16),
    		span: () => (++n).toString(16).padStart(16, '0')
    	};
    	const gateway = new Span('gateway', null, ids, sink);
    	await tracedFetch(gateway, `${base}/rides`, { method: 'POST' });
    	expect(received.at(-1)).toBe(gateway.outgoing());
    	const rides = ridesSpan(
    		new Request(`${base}/rides`, { headers: { traceparent: received.at(-1)! } }),
  3. A check on a busy minute

    One request in a test proves the header is sent. It does not prove you can find one request among many. The checker in section 08 unlocks twelve bikes at once and tries to follow one rider by whatever ids its gateway line carries.

    check-runs.mjs
    const ticketLines = gatewayLog.filter((line) => JSON.stringify(line).includes('"r-88"'));
    // A value is an id only if it is r-88's alone: no other rider's gateway line carries it.
    // Run 1 counted event names such as "unlock_failed" as ids; see run-notes.md.
    const others = gatewayLog.filter((line) => !JSON.stringify(line).includes('"r-88"'));
    const ids = [...new Set(ticketLines.flatMap(idsOn))].filter(
    	(id) => !others.some((line) => JSON.stringify(line).includes(id))
    );
Build UIs?A request’s trace can start in the browser. The error message is where you own it.

Where it already is in your components

If your app uses an error-monitoring SDK, it may already be doing this. Sentry’s browser SDK “reads and further propagates two HTTP headers between your applications”, sentry-trace and baggage, on requests to the origins you list in tracePropagationTargets (Distributed tracing). OpenTelemetry has a fetch instrumentation for the web that does the same with the W3C header. A click in your component becomes the first span of the request.

When you have to own it

The Unlock button is where the rider’s story starts, so it is the best place to start the trace. It makes a trace id, sends it as traceparent, and when the unlock cannot be confirmed it shows the first eight characters as a support reference. The rider reads it out, and support searches for exactly one request. It also does not say “failed” when it does not know: the dock may have released the bike.

A save button that starts a trace in the browser and sends it on the request, with a reference in the error.

ReactAlready in your code
SaveButton.tsx
import { useState } from 'react';

// What a tracing SDK does for every fetch, written out: start a trace in the browser and send
// it on the request, so the server's logs for this click share one id.
function newTraceparent(): { header: string; traceId: string } {
	const hex = (bytes: number) =>
		Array.from(crypto.getRandomValues(new Uint8Array(bytes)), (b) =>
			b.toString(16).padStart(2, '0')
		).join('');
	const traceId = hex(16);
	return { header: `00-${traceId}-${hex(8)}-01`, traceId };
}

export default function SaveButton({ body }: { body: unknown }) {
	const [message, setMessage] = useState('');

	async function save() {
		const { header, traceId } = newTraceparent();
		const response = await fetch('/api/save', {
			method: 'POST',
			headers: { 'content-type': 'application/json', traceparent: header },
			body: JSON.stringify(body)
		});
		setMessage(response.ok ? 'Saved.' : `Not saved. Reference ${traceId.slice(0, 8)}.`);
	}

	return (
		<p>
			<button onClick={save}>Save</button> <span role="status">{message}</span>
		</p>
	);
}

10 / Make the call

Pass the id the moment a request crosses a second process.

Inside one process, a request id on every log line is enough, and a tracing backend would be ceremony. The moment a request crosses into a second process that you want to hear from, the id has to cross with it, and the standard header costs no more than a home-made one.

Revisit it when a new service or vendor joins the path, when support tickets start taking hours to explain, and when a new client, a kiosk or a phone app, starts requests of its own.

Take it with you

Explain it without saying “distributed tracing”: “The first server to see a request gives it an id. Every server after that sends the id along and writes it on every line. When a rider complains, we find their line at the front door and search for its id.” Then open the last feature you built with an AI that calls another service, and check whether anything on the call says which request it was.

Paste into your next prompt, and fill in the blanks

Support must be able to explain any single request from <what a ticket
names, such as a user and a time>, even when many users act in the same second.
- Start a W3C trace context (traceparent) at <the entry point>, or continue
  one if the caller sent a valid header.
- Stamp the trace id on every log line in every service.
- Send traceparent on every outgoing call, including calls to <third parties>.
- Log the trace id beside <the user> at <the entry point>, and show a short
  reference to the user when a request fails.
Add a test that sends a request through <the hops> and asserts every log line
and every outgoing call carries the same trace id.
Connections to follow nextRelated lessons

Take the unlock flow into your editor. Add a notification service that texts the rider, and make its log line findable from the same support reference.

Back to architecture →