A restart pattern is a clue. Trace the check that asked for it.
The incident report says the ticket database slowed down, API replicas restarted, and reconnect work rose during recovery. That sequence deserves attention, but it does not yet prove the probe caused the restarts. A process exit, memory kill, node disruption, or rollout can happen near the same dependency incident.
A process can be running while it initializes, running while a database is unavailable, or running while its work is stuck. A single green/red flag collapses those states. You need to decide which response would help before writing the check.
A probe is a repeated check whose result feeds a controller’s decision. In Kubernetes, the kubelet runs container probes. A failing result is therefore part of the control loop, not just a status message on a dashboard.
Compare a first diagnostic passReveal after making a prediction
- Possible causes
- A dependency-backed liveness check failed; the process crashed or was killed; or a rollout/node event replaced it.
- Discriminating evidence
- Record the initial restart count, inspect Pod events and the previous container's termination reason, then compare timestamps with the configured probe path and dependency errors.
- What this fixture establishes
- The intentionally flawed probe example points liveness at
/readyz, whose contract includes database availability. That configuration predicts the failure path. It is not evidence that any live cluster followed it; the disposable drill must supply those observations.
| Probe | Question | What failure asks Kubernetes to do |
|---|---|---|
| Startup | Has this container finished starting? | Keep readiness and liveness gated. If the startup failure threshold is reached, terminate the container; the restart policy governs restart. |
| Readiness | Should this Pod receive new Service traffic? | Mark it unready after the configured failures. Ordinary Service routing excludes unready endpoints; the container keeps running. |
| Liveness | Has this container entered a state where restarting is warranted? | Terminate it after the configured failures. With a Deployment’s restart policy, it restarts. |
Once startup succeeds, its check stops for that container lifetime. Readiness can change repeatedly. Liveness runs independently of readiness after the startup gate opens; an unready container can still fail liveness and restart. These are the platform’s distinct probe semantics.
- User promise
- A ticket request needs a working database; there is no useful cached response in this example.
- Startup
- Local configuration and an in-memory index take up to 50 seconds to initialize in the fixture.
- Recoverable
- The process can reconnect when a temporary database outage ends.
- Restartable
- A stuck local work loop stops progress and recovers in a fresh process.
An endpoint should answer the question its caller is asking.
Give this service three explicit endpoint contracts. Return a small 200 response for success and 503 for failure. Keep the endpoints cheap, unauthenticated
for the kubelet’s access path, and free of sensitive diagnostics. Avoid redirecting a probe
to a login page.
| Endpoint | Success means | Work to avoid |
|---|---|---|
/startupz | Local initialization has completed. This remains true until this process exits. | Making completion depend on every remote service being available. |
/readyz | The service is initialized, local work can progress, and its required database is usable. | A full user transaction or an unbounded dependency call on every probe. |
/livez | The local execution path being monitored can still make progress. | Database, DNS, or third-party checks that fail in every replica together. |
In a real service, readiness might read a recent, bounded dependency observation from a background monitor. Define how old that observation may be and how recovery refreshes it. A cached “good” result that never expires is a false promise; synchronous probes that pile up slow database calls can become extra load during the outage.
For liveness, choose evidence tied to useful local progress. An HTTP response from the
request event loop can reveal that loop being blocked; a separate health thread returning 200 says little about a deadlocked worker. If a heartbeat drives the check, define when progress
is expected so an idle worker is not mistaken for a stuck one.
Give slow startup its own allowance.
The fixture takes 50 seconds to start. Give startup about 120 seconds of nominal allowance: 24 failed samples at a five-second period. That leaves room for a slow cold start without forcing the running service’s liveness check to wait two minutes before noticing a fault.
This excerpt lives under the container in a Deployment. The named port http maps
to port 8080. The complete manifest and disposable service appear in the drill below.
startupProbe:
httpGet:
path: /startupz
port: http
periodSeconds: 5
timeoutSeconds: 1
failureThreshold: 24
readinessProbe:
httpGet:
path: /readyz
port: http
periodSeconds: 5
timeoutSeconds: 1
failureThreshold: 2
successThreshold: 1
livenessProbe:
httpGet:
path: /livez
port: http
periodSeconds: 5
timeoutSeconds: 1
failureThreshold: 3 timeoutSeconds bounds one attempt. failureThreshold counts
consecutive failed attempts before action. A success breaks that failure streak. Readiness
also uses successThreshold to control recovery; it is one here. Startup and
liveness require a success threshold of one. See the Kubernetes configuration reference.
The two readiness failures tolerate a brief blip before withdrawing traffic. The three liveness failures tolerate a little more before restart. Their nominal windows are about 10 and 15 seconds; they are not precise deadlines measured from the onset of a fault. Probe scheduling, timeouts, and the phase of the first failure affect detection. Termination and endpoint propagation add their own delays.
Choose these numbers from cold-start measurements, normal local pauses, probe latency, and the time users can tolerate a broken replica. CPU throttling or overload can delay a healthy process. Aggressive thresholds can then remove capacity at exactly the wrong time.
Why not just add a long initial delay?Startup gating and steady operation
A fixed delay protects startup only for that fixed duration. A startup probe can succeed early when initialization is quick, while still allowing a bounded slow start. It then gets out of the way. If initialization never succeeds, investigate it instead of increasing the allowance indefinitely.
A readiness check alone does not suppress liveness. If you configure liveness to fail while initialization is still legitimate, the process can be restarted before it ever finishes.
A restart can add a second problem to the first.
Reusing /readyz for liveness looks tidy. For this service, it turns database failure
into a request to restart the application. Nothing in that action makes the database available
again.
livenessProbe:
httpGet:
path: /readyz # Wrong for this service: this includes the database.
port: http
periodSeconds: 5
timeoutSeconds: 1
failureThreshold: 3 Ticket requests cannot complete.
Every replica can see the same remote fault.
Local initialization and connection setup repeat.
Connection churn and lost warm capacity may prolong the outage.
That last step is a risk to investigate, not an automatic numerical claim. This page’s experiment counts restarts in a small model; it does not measure database load or outage duration. It makes the unnecessary intervention visible.
One failure, two restart policies.
Database unavailable from 20 s until 100 s. Both policies start with a ready replica.
Local liveness
- Container
- running
- Restarts so far
- 0
- Liveness failures
- 0 / 3
- Startup failures
- 0 / 24
Database unavailable; leave the recoverable process alive.
Restarts over the full 120 s: 0
Liveness also checks the database
- Container
- starting
- Restarts so far
- 1
- Liveness failures
- 0 / 3
- Startup failures
- 0 / 24
Restart requested. Initialization begins again; traffic stays off.
Restarts over the full 120 s: 2
In the database scenario, local liveness leaves the process running. Readiness withdraws traffic, and the same process can become ready when the dependency returns. The database-dependent policy restarts twice in the authored schedule. In the stuck-process scenario, a restart is useful: both policies recover local progress.
If all three replicas depend on the same failed database, all three may become unready. There is still a user-visible outage. Readiness does not manufacture a working replica; keeping liveness local prevents additional damage while you repair the dependency or activate a designed fallback.
Check the controller’s response, not just the HTTP status.
A health handler test can prove which status code a condition produces. It cannot prove that your manifest points to that handler, the kubelet reaches the port, or traffic stops going to an unready endpoint. Verify the application contract and the deployed reaction separately.
| Introduce | Observe | Expected consequence |
|---|---|---|
| 50-second startup | Startup results, Pod readiness, restart count | Unready during initialization; no premature restart; ready after initialization and readiness success. |
| Database unavailable | Readiness failures, EndpointSlice readiness conditions, restart count | Affected endpoints become unready; restart count stays unchanged. |
| Database restored | Readiness recovery and a ticket request through the Service | Traffic returns without replacing the container. |
| Local progress stuck | Liveness failure events and container restart count | The affected container restarts, passes startup again, and returns to readiness. |
Capture the initial restart count so an earlier crash does not look like a new liveness action. Inspect Pod events and the previous container’s termination reason. An out-of-memory kill, application exit, and liveness-triggered restart need different fixes.
Endpoint readiness is control-plane evidence. Also send representative requests through the Service from another Pod and watch which replica serves them. Direct Pod access and port forwarding bypass ordinary Service selection; they cannot establish that unready replicas have stopped receiving Service traffic.
Run a disposable local drillComplete fixture · Docker, kind, and kubectl
The files live in src/lib/content/lessons/startup-readiness-and-liveness-probes/examples. Work through the commands in that directory one observation at a time. They create an
explicitly named local kind cluster and namespace, load the fixture image, and provide
cleanup.
The fixture uses marker files to simulate database failure and a detected local stall. It does not connect to a real database or deliberately deadlock the event loop. Endpoint behavior and the browser model have automated tests; a live cluster drill remains a separate check.
# Run from this lesson's examples directory. Requires Docker, kind, and kubectl.
# All cluster commands explicitly target a disposable local cluster.
kind create cluster --name probes-lesson
docker build -t probes-lesson:v1 .
kind load docker-image probes-lesson:v1 --name probes-lesson
kubectl --context kind-probes-lesson create namespace probes-lesson
kubectl --context kind-probes-lesson apply -f deployment.yaml
kubectl --context kind-probes-lesson -n probes-lesson rollout status deployment/tickets --timeout=180s
# Observe before injecting a failure. Record restart counts for each Pod.
kubectl --context kind-probes-lesson -n probes-lesson get pods -l app=tickets
kubectl --context kind-probes-lesson -n probes-lesson get endpointslices -l kubernetes.io/service-name=tickets -o yaml
# Start an in-cluster client. Its image is already loaded; this adds no new image.
kubectl --context kind-probes-lesson -n probes-lesson run request-check --image=probes-lesson:v1 --image-pull-policy=Never --restart=Never --command -- node -e 'setInterval(() => {}, 1000)'
kubectl --context kind-probes-lesson -n probes-lesson wait --for=condition=Ready pod/request-check --timeout=60s
# Repeat this request before, during, and after the drill. Record status and replica.
kubectl --context kind-probes-lesson -n probes-lesson exec request-check -- node -e "fetch('http://tickets/tickets').then(async r => console.log(r.status, r.headers.get('x-replica'), await r.text()))"
# A marker simulates an unavailable dependency for ONE replica.
kubectl --context kind-probes-lesson -n probes-lesson exec deployment/tickets -- node -e "require('node:fs').writeFileSync('/tmp/database-down', '')"
# Wait at least 15 seconds, then inspect readiness, restarts, and endpoint conditions.
# No restart is expected. Save the affected Pod name from this output.
kubectl --context kind-probes-lesson -n probes-lesson get pods -l app=tickets
kubectl --context kind-probes-lesson -n probes-lesson get endpointslices -l kubernetes.io/service-name=tickets -o yaml
# Replace AFFECTED_POD with that name; do not let exec choose a different replica.
kubectl --context kind-probes-lesson -n probes-lesson exec AFFECTED_POD -- node -e "require('node:fs').unlinkSync('/tmp/database-down')"
# After recovery, simulate a local progress failure in that same replica.
kubectl --context kind-probes-lesson -n probes-lesson exec AFFECTED_POD -- node -e "require('node:fs').writeFileSync('/tmp/stuck', '')"
kubectl --context kind-probes-lesson -n probes-lesson describe pod AFFECTED_POD
# Expect a liveness failure, one container restart, startup gating, then Ready.
# Marker files live in the container writable layer and disappear on restart.
# Cleanup only this disposable cluster when observations are complete.
kind delete cluster --name probes-lesson
Complete Deployment and Service
apiVersion: apps/v1
kind: Deployment
metadata:
name: tickets
namespace: probes-lesson
spec:
replicas: 3
selector:
matchLabels:
app: tickets
template:
metadata:
labels:
app: tickets
spec:
containers:
- name: tickets
image: probes-lesson:v1 # Build and load the local fixture first.
imagePullPolicy: Never
ports:
- name: http
containerPort: 8080
env:
- name: STARTUP_MS
value: '50000'
resources:
requests:
cpu: 100m
memory: 64Mi
limits:
memory: 128Mi
startupProbe:
httpGet:
path: /startupz
port: http
periodSeconds: 5
timeoutSeconds: 1
failureThreshold: 24
readinessProbe:
httpGet:
path: /readyz
port: http
periodSeconds: 5
timeoutSeconds: 1
failureThreshold: 2
successThreshold: 1
livenessProbe:
httpGet:
path: /livez
port: http
periodSeconds: 5
timeoutSeconds: 1
failureThreshold: 3
---
apiVersion: v1
kind: Service
metadata:
name: tickets
namespace: probes-lesson
spec:
selector:
app: tickets
ports:
- port: 80
targetPort: http
Health endpoint fixture
// Disposable drill fixture. Marker files simulate observations, not a real DB or deadlock.
import { createServer } from 'node:http';
import { existsSync } from 'node:fs';
import { pathToFileURL } from 'node:url';
/**
* @param {string | undefined} path
* @param {{ initialized: boolean, stuck: boolean, databaseDown: boolean }} state
*/
export function statusFor(path, { initialized, stuck, databaseDown }) {
if (path === '/startupz') return initialized ? 200 : 503;
if (path === '/livez') return stuck ? 503 : 200;
if (path === '/readyz' || path === '/tickets') {
return initialized && !stuck && !databaseDown ? 200 : 503;
}
return 404;
}
if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) {
const startedAt = Date.now();
createServer((request, response) => {
const status = statusFor(request.url, {
initialized: Date.now() - startedAt >= Number(process.env.STARTUP_MS ?? 50000),
stuck: existsSync('/tmp/stuck'),
databaseDown: existsSync('/tmp/database-down')
});
response.writeHead(status, {
'Content-Type': 'text/plain',
'Cache-Control': 'no-store',
'X-Replica': process.env.HOSTNAME ?? 'local'
});
response.end(status === 200 ? 'ok\n' : 'unavailable\n');
}).listen(Number(process.env.PORT ?? 8080), '0.0.0.0');
}
Container image
# Learning fixture; pin an approved digest for your own deployment.
FROM node:22-alpine
WORKDIR /app
COPY app.mjs .
USER node
CMD ["node", "app.mjs"]
If the expected withdrawal or restart never happens, check the configured path, named port, probe events, and actual response first. If the response is correct but traffic persists, inspect the routing path and existing connections. If restarts happen without liveness failures, follow the termination reason instead of tuning probe thresholds.
Ask what a fresh process could actually repair.
The database is unavailable. What should these replicas report?
Ticket requests require the database. The application can reconnect without restarting, and its local work loop is still responsive.
A liveness probe earns its place when it can detect a failure that a restart is likely to repair. If your process already exits reliably on unrecoverable local failure, the restart policy may be enough. Adding a second failure detector is a design decision with false positives and operating costs.
Keep monitoring alongside probes. Alert on user-visible errors and latency, sustained loss of ready capacity, and repeated restarts. A green health endpoint does not establish that real requests meet their promise.
Leave a short note with the configuration: what each endpoint means, where its signal comes from, why its tolerances fit, and how to exercise its failure and recovery. That note helps the next person resist “simplifying” three different questions into one.
The question to keep: if this check fails across every replica, does the action help the system recover?
Connections to follow nextRelated lessons
Retry, backoff, and idempotency explains why repeated attempts need a budget. Restarts can create a similar surge of repeated connection work.
Backpressure and queues examines what to do when incoming work exceeds capacity. Restarting overloaded workers may remove capacity instead of relieving the bottleneck.
Timeouts, deadlines, and races separates one attempt’s time limit from the whole operation’s budget.
Sources and scopePrimary documentation · checked September 27, 2026
- Kubernetes: Liveness, Readiness, and Startup Probes — probe meaning, gating, and when a liveness probe is useful.
- Kubernetes: Configure Liveness, Readiness and Startup Probes — thresholds, timeouts, and the warning about poorly implemented liveness checks.
- The browser timeline is a tested teaching model. The local fixture checks endpoint contracts. Neither is presented as a production incident recording or proof of a particular cluster’s timing.