01 / The prompt
“Build the unlock flow, and log each request.”
A city bike-share: a rider taps Unlock, a gateway calls the rides service, rides puts a $1 hold on the card through payments and asks the dock controller to release the bike. You ask an agent for it, with logging. What comes back works, and every service writes neat JSON lines.
Then a ticket arrives: “r-88, about 08:14, bike B-12. The app said something went wrong, and there’s a $1 charge on my card.” The gateway log has one line for r-88: a 502 after two seconds. Rides, payments, and the dock each have lines from 08:14, and eleven other riders were unlocking bikes in the same minute. Which lines are r-88’s?
The prompt said to log each request. It did not say that a request is one thing across four programs, so each program logged its own. That is the question it never answered: how does one request keep one identity across every hop?
02 / Name the shape
One id, every hop.
Observability across a request means one request can be followed through every service it touches, after the fact, from whatever the user can tell you. The mechanism is context propagation: the first service gives the request an id, and every service passes it on, on every call, and writes it on every line.
A request’s identity is created once, at the edge, and inherited everywhere else. No service invents its own.
Here is who knows what about r-88’s unlock without it.
| Service | Knows | Cannot tell you |
|---|---|---|
| Gateway | The rider, the bike, the status, how long it took | What happened inside. |
| Rides | The rider and bike, the hold, each dock attempt | What the dock did after it stopped waiting. |
| Payments | An amount and a card | Which unlock the hold was for. |
| Dock controller | A bike and a station | Who asked, or whether anyone was still listening. |
Words to put in a prompt or a review
- Trace
- Everything one request did, across every service, under one id.
- Span
- One service’s part of it, with its own id and its caller’s id as its parent.
- Context propagation
- Passing the trace id and the current span id on every call.
- traceparent
- The W3C header that carries them:
00-traceid-spanid-flags. - Correlation
- Joining lines by a shared id, not by time or by guessing from fields.
- Support reference
- A short piece of the trace id shown to the user when something fails.
Logs, metrics, and tracesWhich one answers which question
Metrics say how often and how slow, for everyone: a rising 502 rate at 08:14. Logs say what one service did. Traces say what one request did everywhere. The trace id is what lets a log line and a trace point at the same request; tools like OpenTelemetry record spans and attach the trace id to logs for that reason. You can start with logs alone, as this lesson does: the propagated id is the part that matters, and a tracing backend can come later.
03 / Follow one unlock
Twelve riders, one slow dock, and a ticket.
First the minute as it was logged with service-local ids, and support searching it by time. Then the same minute with a trace id passed on every call, and support searching by that id. Open Try it to search for any rider yourself.
Which of these lines are r-88’s?
Service-local request ids · 08:14:00.000
gateway.log
- No lines yet
rides.log
- No lines yet
payments.log
- No lines yet
dock.log
- No lines yet
Replaying the logs as they were written
Twelve riders tap Unlock, 150 ms apart.
Each service writes its own log line with its own request id. Nothing ties one service’s id to another’s.
Reduced motion: choose a scene to see its completed state.
Read this scene
Each service writes its own log line with its own request id. Nothing ties one service’s id to another’s.
Service-local ids, 0 ms after 08:14:00. gateway: nothing.rides: nothing.payments: nothing.dock: nothing.
Watch restarts the story when you come back. Step through keeps your step. Try it lets you search the same minute as support.
04 / Read the shape
A header, a span, and a call that passes it on.
Basic form is the header format, and what to do with a bad one. In the wild is the span each service opens: continue the caller’s trace, or start one. At the call site the context crosses the wire, and every outgoing call carries it.
The W3C traceparent header: version, a 32-hex trace id, the caller’s 16-hex span id, and flags. Anything malformed, uppercase, or all zeros is no context, and the service starts a new trace.
/** One hop's view of a request: which trace it belongs to and which span called it. */
export type TraceContext = { traceId: string; spanId: string; sampled: boolean };
const TRACEPARENT = /^([0-9a-f]{2})-([0-9a-f]{32})-([0-9a-f]{16})-([0-9a-f]{2})$/;
/** Read a `traceparent` header. Anything malformed, or all zeros, is no context at all. */
export function parseTraceparent(header: string | null | undefined): TraceContext | null {
const match = header ? TRACEPARENT.exec(header.trim()) : null;
if (!match) return null;
const [, version, traceId, spanId, flags] = match;
if (version === 'ff' || /^0+$/.test(traceId) || /^0+$/.test(spanId)) return null;
return { traceId, spanId, sampled: (parseInt(flags, 16) & 1) === 1 };
}
export function formatTraceparent(context: TraceContext): string {
return `00-${context.traceId}-${context.spanId}-${context.sampled ? '01' : '00'}`;
} // TraceContext is one hop's view of a request: which trace it belongs to and which span called it.
type TraceContext struct {
TraceID string `json:"traceId"`
SpanID string `json:"spanId"`
Sampled bool `json:"sampled"`
}
var traceparent = regexp.MustCompile(`^([0-9a-f]{2})-([0-9a-f]{32})-([0-9a-f]{16})-([0-9a-f]{2})$`)
// ParseTraceparent reads a traceparent header. Anything malformed, or all zeros, is no context.
func ParseTraceparent(header string) *TraceContext {
m := traceparent.FindStringSubmatch(strings.TrimSpace(header))
if m == nil || m[1] == "ff" || strings.Trim(m[2], "0") == "" || strings.Trim(m[3], "0") == "" {
return nil
}
flags, _ := strconv.ParseUint(m[4], 16, 8)
return &TraceContext{TraceID: m[2], SpanID: m[3], Sampled: flags&1 == 1}
}
func FormatTraceparent(c TraceContext) string {
flags := "00"
if c.Sampled {
flags = "01"
}
return "00-" + c.TraceID + "-" + c.SpanID + "-" + flags
} The behavior these examples promiseChecked by 38 shared cases in TypeScript and Go
- A traceparent is two lowercase hex digits, 32, 16, and 2, joined by dashes, with
surrounding spaces ignored. Version
ff, an all-zero trace id, or an all-zero span id makes it invalid. The sampled flag is the lowest bit. - A span continues a valid incoming trace, with the caller’s span as its parent, or starts a new trace with no parent. Its outgoing header names itself as the parent.
- The simulated minute: twelve riders 150 ms apart; gateway to rides 5 ms; a hold after 80 ms; station 3’s dock answers in 120 ms and station 7’s in 3000; rides waits 1000 ms per attempt, twice, then answers 502 and keeps the hold; the dock logs its unlocks when they finish.
- Support starts from the gateway line for the rider. With a trace id it takes every line in that trace; without one, every line written while the request was open.
Every expectation in the shared cases was produced by a separate model written from these
rules, kept beside the examples in examples/model/, not copied from either
implementation: twelve headers, both minutes line for line, and what support finds for
each of the twelve riders in each shape.
Reading the TypeScriptA regex, a class, and Headers
One regular expression checks the whole header, so a header is either a context or null, never half of one. Span keeps the service and log sink
in private fields; outgoing() is the only way to get a header out of it. new Headers(init.headers) in tracedFetch keeps whatever headers
the caller set and adds one.
Reading the GoA pointer for no context, and a handler
ParseTraceparent returns *TraceContext, nil for no context. strings.Trim(id, "0") == "" is the all-zero check. The rides handler reads
the header from r.Header and continues the trace before doing anything else.
Run it yourselfNo dependencies
Copy the complete TypeScript file and run node --experimental-strip-types trace.ts with Node 22.18 or later. For Go, save main.go next to this go.mod and run go run .. Both print:
module heyrian.dev/lessons/observability-across-a-request
go 1.23
local: r-88 got 502 after 2090 ms; found gateway 5, rides 11, payments 4, dock 4; the bike's unlock missed traced: r-88 got 502 after 2090 ms; found gateway 1, rides 4, payments 1, dock 2; the bike's unlock found
The minute is simulated, and the ids count up so both languages print the same thing. Real services use random ids.
05 / Review the agent’s diff
“Adds request ids to the rides service.”
Structured logs with ids are a real improvement over printed strings. Read where this id comes from, and where it goes.
06 / How it fails
A trace breaks at the first hop that forgets it.
Each way the unlock can go wrong, what the rider sees, and what the traced build lets support see. The first row comes from the shared cases; the rest are authored from the same example.
| What goes wrong | What the rider sees | What support finds with a trace | Without one |
|---|---|---|---|
| Slow dock, unlock after the rider gave up | An error, a hold, and a bike that is free. | 8 lines, including the dock unlocking B-12 a second after the 502. | 24 lines, most of them other riders’; the late unlock missed |
| Payments down | An error, no hold. | The rides span, a failed payments call, and no dock call. | Payments’ own error, with no rider on it |
| A hop drops the header | Nothing different. | A trace that stops at that hop; the next service starts a new trace. | No difference to notice |
| Two taps, two requests | One unlock, maybe two holds. | Two traces, each with its own hold. | Two holds and no way to tell which tap made which |
| A malformed or all-zero header arrives | Nothing. | A new trace from that hop; the bad header is not trusted. | Not applicable |
| A trace id in the user-facing error | A reference to read out. | The whole request from that reference. | A time and a guess |
The slow dock is also a failure-mode question: an unlock whose outcome is unknown should not be reported as failed, and the hold should not outlive it. Tracing does not fix that; it is how you find out it happened. Keeping one slow dependency from slowing everything else is Containing failure.
07 / Is it worth it?
You pay one header and one field per line. Here is what they buy.
Service-local logs are simpler to write and are enough while one service does all the work. Hold both shapes up against the changes a system like this gets.
| Change | Service-local ids | Propagated context |
|---|---|---|
| A second client: the kiosk at each station | Another source of lines to line up by time. | The kiosk starts or sends a traceparent; its requests are traces like any other. |
| Replace the dock vendor | New log formats to learn to join. | The new vendor receives the header. Whether it logs it is a question for the contract. |
| Change a rule: three dock attempts, not two | One more line to find by time. | One more span under the same trace. |
| A second team owns payments | Their logs never meet yours. | A shared id is the only shared language the two teams’ logs need. |
Before you add propagation, decide what you would measure and what you would accept:
- Time to explain a ticket, from ticket to cause, for the next ten tickets. Take it before the change, and after.
- Share of log lines with a trace id, per service. Anything under all of them is a hop that breaks the chain.
- Traces with one span where you expected several: a hop that dropped the header, found by counting instead of by a bad day.
This page did not run a bike-share, so it has no numbers of that kind to give you. Which requests count as failures in the first place is Defining success.
08 / Ask for it
Two prompts, twelve riders, one ticket to follow.
We sent two agents the same request for this unlock flow at the same time, both running Claude Sonnet. Both prompts asked for JSON log lines in each service. One added a paragraph: support must take a ticket that names a rider and a time, find that one request, and find every line it left and every call it made, with one search, even when many riders unlock in the same second. It named no header and no standard. Then a script ran each build against fake payments and dock services, unlocked twelve bikes 150 ms apart with one slow dock, and followed whatever id r-88’s gateway line carried.
| What the checker looked at | Plain prompt | Architecture prompt |
|---|---|---|
| r-88’s unlock, in a minute of twelve | 502 after 2.1 s | 502 after 2.1 s |
| Ids on r-88’s gateway log line | None | 1 |
| Rides log lines linked by those ids | 0 of 37 (3 mention r-88) | 8 of 75 (8 mention r-88) |
| Calls to payments that carry them | 0 of 12 | 1 of 12 |
| Calls to the dock that carry them | 0 of 13 (2 were for B-12) | 2 of 13 (2 were for B-12) |
| Header that carried them | None | x-request-id |
The plain build logged carefully and joined nothing. Its rides lines carry the rider, bike, station, and hold id, so a patient person could match them to the gateway by rider. But its gateway line for r-88 carries no id, and neither the payments call nor the two dock calls for B-12 carry anything that says which request they belong to.
The architecture build did what this lesson teaches, under another name. It made a UUID at
the gateway, wrote it on every line with the rider, and sent it as X-Request-Id on
every call, including both dock attempts. One search for that id found exactly r-88’s lines and
calls.
log(RIDES_LOG, { event: "unlock_requested", rider, bike, station });
const hold = await placeHold(rider);
if (!hold) {
log(RIDES_LOG, { event: "hold_failed", rider, bike, station });
sendJson(res, 502, { unlocked: false, error: "payment hold failed" });
return;
}
log(RIDES_LOG, { event: "hold_placed", rider, bike, station, holdId: hold.holdId });
const released = await releaseBike(bike, station); const { status, body } = await postJson(
`${paymentsUrl}/holds`,
{ rider: ctx.rider, amountCents: 100 },
{ "X-Request-Id": ctx.requestId },
); Two details are worth the prompt line. The header was home-made, so a vendor, a proxy, or a
tracing library would not recognize it; the W3C traceparent is the one they
already pass on. And the id starts at the gateway, so the rider’s screen has nothing to read
out to support. The line the runs point to: start a W3C trace context in the client, continue it in every service, and show the user
a short reference when a request fails.
How the runs were made and checkedOne run each, recorded as written
- Both agents received the prompts word for word, in fresh contexts, in the same message. Neither was told about the other, this lesson, or the checker. Besides the Architecture paragraph and the output folder, the prompts differed in one line: the ports each could use to test with, so their stand-ins could not collide.
- The files each agent wrote are kept byte for byte, with checksums, beside this lesson’s examples. The checker restores each build, starts it with fake payments and dock services that record every call, unlocks twelve bikes 150 ms apart with one slow dock, and follows any id on r-88’s gateway line.
- This is one sample of each prompt, not a measurement of a model. The transcript audit and every number are in the run notes beside the examples.
09 / Hold it there
Make a missing trace id a failing test.
Propagation breaks quietly: a new call written with a bare fetch, a new service
that logs with console.log. Three checks catch it.
The standard, not a home-made header
Use the W3C Trace Context format. Its
traceparentisversion-trace-id-parent-id-trace-flags, an all-zero trace id “is considered an invalid value”, and when the header is invalid “the vendor creates a newtraceparentheader” (W3C Trace Context). Tracing libraries, proxies, and vendors already speak it, so a hop you do not own may pass it on for you.A test that follows one request
Send one request through the hops you own, with the calls to payments and the dock pointed at a local server that records headers, and assert every log line and every outgoing call carries the same trace id. This lesson’s suite does that for the gateway-to-rides hop.
trace.spec.ts it('sends the current span as the parent, and the callee continues the trace', async () => { const sink: LogLine[] = []; let n = 0; const ids: Ids = { trace: () => 'cd'.repeat(16), span: () => (++n).toString(16).padStart(16, '0') }; const gateway = new Span('gateway', null, ids, sink); await tracedFetch(gateway, `${base}/rides`, { method: 'POST' }); expect(received.at(-1)).toBe(gateway.outgoing()); const rides = ridesSpan( new Request(`${base}/rides`, { headers: { traceparent: received.at(-1)! } }),A check on a busy minute
One request in a test proves the header is sent. It does not prove you can find one request among many. The checker in section 08 unlocks twelve bikes at once and tries to follow one rider by whatever ids its gateway line carries.
check-runs.mjs const ticketLines = gatewayLog.filter((line) => JSON.stringify(line).includes('"r-88"')); // A value is an id only if it is r-88's alone: no other rider's gateway line carries it. // Run 1 counted event names such as "unlock_failed" as ids; see run-notes.md. const others = gatewayLog.filter((line) => !JSON.stringify(line).includes('"r-88"')); const ids = [...new Set(ticketLines.flatMap(idsOn))].filter( (id) => !others.some((line) => JSON.stringify(line).includes(id)) );
Build UIs?A request’s trace can start in the browser. The error message is where you own it.
Where it already is in your components
If your app uses an error-monitoring SDK, it may already be doing this. Sentry’s browser
SDK “reads and further propagates two HTTP headers between your applications”, sentry-trace and baggage, on requests to the origins you list in tracePropagationTargets (Distributed tracing). OpenTelemetry has a fetch instrumentation for the web that does the same with the W3C
header. A click in your component becomes the first span of the request.
When you have to own it
The Unlock button is where the rider’s story starts, so it is the best place to start the
trace. It makes a trace id, sends it as traceparent, and when the unlock
cannot be confirmed it shows the first eight characters as a support reference. The rider
reads it out, and support searches for exactly one request. It also does not say “failed”
when it does not know: the dock may have released the bike.
A save button that starts a trace in the browser and sends it on the request, with a reference in the error.
import { useState } from 'react';
// What a tracing SDK does for every fetch, written out: start a trace in the browser and send
// it on the request, so the server's logs for this click share one id.
function newTraceparent(): { header: string; traceId: string } {
const hex = (bytes: number) =>
Array.from(crypto.getRandomValues(new Uint8Array(bytes)), (b) =>
b.toString(16).padStart(2, '0')
).join('');
const traceId = hex(16);
return { header: `00-${traceId}-${hex(8)}-01`, traceId };
}
export default function SaveButton({ body }: { body: unknown }) {
const [message, setMessage] = useState('');
async function save() {
const { header, traceId } = newTraceparent();
const response = await fetch('/api/save', {
method: 'POST',
headers: { 'content-type': 'application/json', traceparent: header },
body: JSON.stringify(body)
});
setMessage(response.ok ? 'Saved.' : `Not saved. Reference ${traceId.slice(0, 8)}.`);
}
return (
<p>
<button onClick={save}>Save</button> <span role="status">{message}</span>
</p>
);
}
10 / Make the call
Pass the id the moment a request crosses a second process.
Inside one process, a request id on every log line is enough, and a tracing backend would be ceremony. The moment a request crosses into a second process that you want to hear from, the id has to cross with it, and the standard header costs no more than a home-made one.
Revisit it when a new service or vendor joins the path, when support tickets start taking hours to explain, and when a new client, a kiosk or a phone app, starts requests of its own.
Take it with you
Explain it without saying “distributed tracing”: “The first server to see a request gives it an id. Every server after that sends the id along and writes it on every line. When a rider complains, we find their line at the front door and search for its id.” Then open the last feature you built with an AI that calls another service, and check whether anything on the call says which request it was.
Paste into your next prompt, and fill in the blanks
Support must be able to explain any single request from <what a ticket names, such as a user and a time>, even when many users act in the same second. - Start a W3C trace context (traceparent) at <the entry point>, or continue one if the caller sent a valid header. - Stamp the trace id on every log line in every service. - Send traceparent on every outgoing call, including calls to <third parties>. - Log the trace id beside <the user> at <the entry point>, and show a short reference to the user when a request fails. Add a test that sends a request through <the hops> and asserts every log line and every outgoing call carries the same trace id.
Connections to follow nextRelated lessons
- Defining success says when something went wrong; this lesson finds which request it was.
- Thinking in failure modes would have asked what the dock does after rides stops waiting.
- Containing failure keeps the slow dock from holding every unlock.
- Errors across a boundary covers what an error should carry when it crosses from one service to the next.