“The rollout was green” and “a few requests failed” can both be true.
Suppose support reports several checkout requests ending with connection resets during a rollout. The Deployment became available, error rates returned to baseline, and the team has no confirmed duplicate orders. This is a useful starting report; it is not yet evidence that the process exited too early or that Kubernetes sent traffic to a terminating Pod.
Keep the questions separate: which requests failed, which Pod handled them, whether the process was terminating, whether that request had already been accepted, and what the client did next. A retry may succeed while the first attempt is still completing, which makes idempotency part of the investigation.
Compare a first diagnostic passReveal after making a prediction
- Competing explanations
- The process exited while work was active; new traffic reached a Pod during route convergence; a handler or dependency timed out; or the client abandoned and retried the request.
- Useful evidence
- Request IDs and durations, response/reset type, Pod UID, termination start and exit time, application signal logs, EndpointSlice conditions, and client retry records.
- What the lesson establishes
- The sample configuration defines a lifecycle to investigate. It does not claim a real cluster had these timings or that every client follows the same retry policy.
Withdrawal and process shutdown overlap in time.
When a Pod is deleted, Kubernetes starts its termination grace countdown and begins
removing the Pod from matching EndpointSlices while running any preStop hook.
The hook runs synchronously; after it completes, the container runtime sends the stop
signal (commonly SIGTERM). At grace expiry, remaining processes can be
forcefully terminated.
EndpointSlice updates and application shutdown are concurrent parts of the transition. Endpoint changes need to propagate through consumers and proxies; they do not form a guarantee that all traffic has stopped before the process receives its signal. This is why the handler should stop accepting new work promptly and remain robust to requests that arrive during convergence.
“Drain” has multiple meanings. Stop accepting new requests on a listener; let accepted requests finish; close or wait for keep-alive connections; separately handle upgraded or long-lived sessions. A web server’s graceful close may cover ordinary requests while WebSockets, streams, or external load balancers need their own protocol.
Choose what happens to new work and active work.
| Request state | Shutdown behavior | Evidence to collect |
|---|---|---|
| Not accepted yet | Mark the instance unready and stop admitting work when shutdown begins. During routing convergence, return a deliberate retryable response or close safely. | Request arrival time, Pod UID, readiness and endpoint condition, response status. |
| Accepted and in flight | Allow bounded work to finish. At the deadline, cancel or close according to the application contract. | Request duration, completion/cancellation reason, drain deadline, client outcome. |
| Long-lived or upgraded | Define a separate reconnect, close-frame, or migration policy; ordinary HTTP server close behavior may not cover it. | Connection type, session age, close code, reconnect and duplicate-action behavior. |
Readiness indicates whether an instance should receive new work under the service’s routing contract. It is not a command that instantly empties every connection. A shutdown flag should be idempotent: Kubernetes lifecycle hooks can be retried, and repeated drain requests must not reopen service or extend the deadline without limit.
Give the application room to finish, with a hard stop.
The example allows five seconds for the hook and up to twenty-five seconds for application shutdown inside a forty-second Pod grace period. Ten seconds remain for scheduling, signal delivery, and other termination overhead. Those are teaching values, not universal defaults: derive them from measured request durations, client deadlines, and the reliability cost of waiting.
If requests can validly take longer than the drain window, either revise the request contract or accept that some work will be cut off. Infinite waiting can strand capacity and block rollout progress. A deadline is a deliberate boundary, and its expiration should be observable.
apiVersion: apps/v1
kind: Deployment
metadata:
name: reports-api
spec:
replicas: 3
selector:
matchLabels:
app: reports-api
template:
metadata:
labels:
app: reports-api
spec:
# 5s preStop allowance + up to 25s active-request drain + 10s margin.
terminationGracePeriodSeconds: 40
containers:
- name: api
image: reports-api:reviewed-release
ports:
- name: http
containerPort: 8080
lifecycle:
preStop:
httpGet:
path: /drain
port: http
readinessProbe:
httpGet:
path: /readyz
port: http
periodSeconds: 3
timeoutSeconds: 1
failureThreshold: 1
Make shutdown explicit in the server and bounded by a deadline.
Both examples set a drain state used by readiness, then stop accepting new connections and
wait for ordinary in-flight work. The TypeScript version forces remaining connections
closed after its deadline. The Go version passes a deadline context to http.Server.Shutdown and calls Close if graceful shutdown expires. Neither example automatically solves
WebSocket or external load-balancer draining.
import { createServer } from 'node:http';
import { setTimeout as delay } from 'node:timers/promises';
let draining = false;
const server = createServer(async (request, response) => {
if (request.url === '/readyz') {
response.writeHead(draining ? 503 : 200).end();
return;
}
if (request.url === '/drain') {
draining = true;
await delay(5_000); // bounded routing-propagation allowance for this example
response.writeHead(200).end('draining');
return;
}
if (draining) {
response.writeHead(503, { connection: 'close' }).end('instance is draining');
return;
}
await delay(3_000); // representative in-flight request
response.writeHead(200).end('request complete');
});
server.listen(8080);
process.once('SIGTERM', () => {
const forceStop = setTimeout(() => server.closeAllConnections(), 25_000);
forceStop.unref();
server.close(() => {
clearTimeout(forceStop);
process.exitCode = 0;
});
});
package main
import (
"context"
"log"
"net/http"
"os"
"os/signal"
"sync/atomic"
"syscall"
"time"
)
func main() {
var draining atomic.Bool
mux := http.NewServeMux()
mux.HandleFunc("GET /readyz", func(w http.ResponseWriter, r *http.Request) {
if draining.Load() {
http.Error(w, "draining", http.StatusServiceUnavailable)
return
}
w.WriteHeader(http.StatusOK)
})
mux.HandleFunc("GET /drain", func(w http.ResponseWriter, r *http.Request) {
draining.Store(true)
time.Sleep(5 * time.Second) // bounded routing-propagation allowance for this example
w.WriteHeader(http.StatusOK)
})
mux.HandleFunc("GET /work", func(w http.ResponseWriter, r *http.Request) {
if draining.Load() {
http.Error(w, "instance is draining", http.StatusServiceUnavailable)
return
}
select {
case <-time.After(3 * time.Second):
w.Write([]byte("request complete"))
case <-r.Context().Done():
return
}
})
server := &http.Server{Addr: ":8080", Handler: mux}
serveErr := make(chan error, 1)
go func() { serveErr <- server.ListenAndServe() }()
signals := make(chan os.Signal, 1)
signal.Notify(signals, syscall.SIGTERM, os.Interrupt)
select {
case <-signals:
case err := <-serveErr:
if err != nil && err != http.ErrServerClosed {
log.Fatal(err)
}
return
}
draining.Store(true)
ctx, cancel := context.WithTimeout(context.Background(), 25*time.Second)
defer cancel()
if err := server.Shutdown(ctx); err != nil {
log.Printf("grace period expired; close remaining connections: %v", err)
_ = server.Close()
}
if err := <-serveErr; err != nil && err != http.ErrServerClosed {
log.Fatal(err)
}
}
Exercise the transition and follow one request end to end.
In a disposable environment, run a request that lasts longer than a few seconds while deleting its Pod, then repeat with a request near the configured deadline. Record whether it completed, was cancelled, or was retried. Also issue fresh requests during the transition. Repeat with a persistent connection if the service uses them.
Compare application timestamps to Pod events and EndpointSlice state, then verify the client-visible result. A port-forward-only test skips much of the Service routing path; it can validate the process behavior but not the whole traffic transition. Observe failed and successful requests so a temporary error spike is not hidden by the final rollout status.
Graceful means predictable completion or predictable cancellation.
Before shipping a shutdown change, name the request classes the service promises to finish, the deadline for each, the client retry and idempotency behavior, and the evidence that demonstrates the result. If the measured failures persist, keep the cause open: routing propagation, application close behavior, load-balancer policy, and client retries remain distinct places to investigate.
Follow the remaining resets
A rollout still shows occasional resets after the drain handler is added. What is the most useful next move?
Operational rule to carry forward: stop admitting new work, finish accepted work only within a measured budget, and verify what users and clients actually observe across the complete request path.
References: Kubernetes container lifecycle hooks, Pod lifecycle, Node.js HTTP server close, and Go HTTP server shutdown. Accessed 2026-10-01.