← DevOps
Technique Incident response and operations

Triage and evidence
gathering

The first plausible cause is a hypothesis, not a finding.

At 14:10, the checkout dashboard shows more failed requests. A deployment finished at 14:04. The on-call channel says, “The release broke payments.” You have two useful facts—an increase and a nearby release—and one untested explanation. Your first job is to learn how many users are affected and what observation would change your next decision.

The judgment to keep

State what you know, what you infer, and what you will check next. Reduce user impact when needed while keeping the investigation honest.

Incident response · empirical troubleshooting · handoff
01 / Scope the report

“Errors are up” needs a population, a window, and a comparison.

To practice, use this illustrative dashboard snapshot: 68 of 1,000 checkout requests failed in the five minutes after the alert (6.8%). In the comparable five-minute window before it, 2 of 1,000 failed (0.2%). The deployment completed between those windows. These figures are authored teaching data; they are not observations from a live service.

Ask which users, routes, regions, and request types are affected. Is latency rising along with failures? Are attempts counted once, or can client retries inflate the denominator? Does the failure prevent purchase, or is a noncritical status panel slow? Impact determines whether to mitigate immediately and what comparison is meaningful.

Write down three lines: observed fact, current hypothesis, and the next question. Keep the words distinct even when the chat room is moving quickly.
Compare a precise first noteReveal after writing your own
Incident note / 14:10 UTCReport the observation without assigning cause.
Observed
In this teaching snapshot, 68/1,000 checkout requests failed in the alert window, versus 2/1,000 in the prior comparable window.
Known context
A deployment completed at 14:04. The time ordering is known; causation is not.
Unknown
Whether failures cluster by version, region, route, dependency, or client; and whether retries affect the count.
Next question
Do failures correlate with the new version after controlling for route and region, and do dependency signals show a matching change?
02 / Build the timeline

Put events in order before you connect them.

Align the alert, first failing request, deployment start and completion, dependency alarms, traffic shifts, and any operator changes. Normalize timestamps and time zones. Preserve request IDs, release identifiers, and the original metric window so later comparisons refer to the same events.

Timelines expose gaps: an error could precede the rollout; a dependency slowdown may begin first; or the failure may appear only as traffic shifts to new instances. A timeline supports a causal story, but adjacency alone does not establish one.

Incident timeline / example data
TimeRecordWhat it establishes
13:59–14:04Illustrative baseline window, 2/1,000 failedComparison rate for this exercise; validate equivalent traffic and counting.
14:04Deployment marked completeController reported completion; not proof every instance behaves correctly.
14:05–14:10Illustrative alert window, 68/1,000 failedFailure rate rose in the chosen window; cause remains open.
14:10On-call receives alertDetection time, which may lag the onset.
03 / Keep alternatives open

A useful hypothesis predicts what you should observe.

Do not collect evidence without a question. Write down a few explanations that account for the same report, then state what each predicts. Choose checks that could make one explanation less likely as well as more likely.

Working hypothesis register
Possible explanationPredictionEvidence that would change belief
New application version has a regressionFailures are more common on the new version for comparable routes and regions.Version-tagged request outcome and error type, with traffic share and denominator.
Payment dependency is degradedFailures cross application versions and align with dependency timeouts or saturation.Dependency latency/error signals and request traces spanning the call.
Traffic or configuration changedFailures cluster after a routing, secret, flag, or traffic-shift event.Change records, instance configuration identity, and results by route/region.
04 / Choose evidence

Ask for the smallest safe check that separates explanations.

A version comparison by route may tell you whether the new build is implicated; dependency traces may show a shared downstream failure. Check the denominator and traffic mix before interpreting either. A chart aggregated across versions can hide a clear split, while a few handpicked traces can exaggerate one.

Good troubleshooting names the question, the expected observations, the owner, and the next decision. Prefer read-only queries and narrowly scoped comparisons first. Preserve raw events before dashboards roll up or logs expire. Change one variable at a time when the service is stable enough; during a severe incident, choose the fastest safe mitigation and document the tradeoff.

Evidence request“Is this tied to the new version?”
Compare
Failure rate and error class by application version, route, and region over the same interval.
Control
Check request counts, rollout traffic share, and client retry accounting before comparing percentages.
Decision it informs
Whether to pause or reverse rollout, investigate a shared dependency, or continue to another hypothesis.
05 / Reduce user impact

Mitigation can be justified before root cause is certain.

Empirical thinking does not mean waiting for perfect proof while users are harmed. It means choosing an action for an explicit reason, considering its downside, and checking whether the outcome matches the prediction. Pausing rollout can stop exposure to a suspected version. Rolling back may restore service, but can also remove a needed fix or leave incompatible data changes. A feature flag, traffic shift, or dependency failover may be safer in a particular system.

Before acting, say what impact warrants the change, how to reverse it, which signal will show improvement, and who is watching. Record the pre-change state and decision time. Afterward, verify user-facing outcomes; “the dashboard turned green” is not enough if the affected requests still fail.

Try a decision

Choose a response under uncertainty

The illustrative error increase began just after the deployment. Checkout is still failing for users. What is the best next move?

Your next move
06 / Hand off what you know

A good handoff preserves uncertainty and makes the next step obvious.

Keep a short incident record with impact and window, verified observations, current hypotheses, actions and their results, remaining risks, evidence links, and a named next owner. Separate mitigation from the later root-cause analysis. Once service is stable, test the causal explanation against the timeline and look for evidence that could disconfirm it.

For the example: the failure rate increased in the post-deployment window; the deployment is a candidate cause; no version-sliced evidence has yet established that link. The next step is to compare outcomes by version, route, and region while checking dependency signals. If impact worsens before that check completes, act on the agreed mitigation threshold and continue gathering evidence.

References: Google SRE, Effective Troubleshooting and Managing Incidents; Google’s Incident Management Guide. Accessed 2026-10-01.