Two populations during change
The second part of the lab holds a request on A using a script-controlled event. While that request is pending, new configuration routes subsequent requests to B. The master receives HUP and the fixture waits for the new generation. A new request reaches B successfully, but the first is still associated with A. This provides a concrete observation of different populations during change. Success for new clients does not, by itself, determine outcomes for clients that were already connected.
The deadline belongs to the worker
The initial file configures worker_shutdown_timeout=1s. The short value makes behavior observable in the lab; it is not a production recommendation. Once the old worker begins shutdown, this deadline limits waiting before an attempt to close open connections. In the recorded run, the held request’s client receives RemoteDisconnected while its backend still waits for the event. The fixture disables upstream retries and uses a longer proxy_read_timeout. Evidence distinguishes the shutdown deadline from replay or normal completion of the request.
Losing a response does not define outcome
After client failure, the script confirms that the backend started work and has not completed it. Only then does it release the event and observe completion. The handler was deliberately written to continue independently of its socket; this is not a claim about every application. The experiment demonstrates that cancellation cannot be inferred from a transport error. For a fictional instruction with a business effect, status must be checked by identifier and the agreed cancellation, deduplication or reconciliation contract must be followed.
Choose deadlines with evidence
Before retiring instances in a real environment, characterize durations, long-running requests, remaining capacity and maintenance-window constraints. Also identify who can cancel work, what becomes durable and how an outcome is checked after losing the response. Increasing every deadline can extend occupancy and maintenance; shortening them can increase interruptions. The decision should make those trade-offs explicit and include change abort criteria. The lab provides a reproducible mechanism, but it neither measures representative load nor chooses values for your service.
Run the go/no-go meeting
Use the worksheet below in a fictional meeting between the project manager, APS and development. One person presents new-traffic evidence, another explains old requests and another validates state continuity. Fill in limits, owners, recovery and the next checkpoint using exercise data or an authorized environment. The objective is to practice decisions, not fill every field with “OK”. If the two uncertain outcomes have no owner, RUN handover is incomplete. The worksheet is a proposed exercise; no human session has been performed to validate it.
Communicate and close without omissions
An English update can say “New traffic is healthy; two outcomes remain under reconciliation with named owners”, followed by permitted identifiers, owners and the next update deadline. If rollback occurs, confirm the generation actually serving traffic and continue reconciling earlier work: restoring routing does not reverse application effects. Retain enough evidence for the next team and protect session data. At closure, distinguish what ran in the lab, what was merely proposed and what still needs representative rehearsal and independent review.
DR proposed change rehearsal: affinity and retirement
No human workshop has been performed.
Fictional example; not a BNP Paribas procedure.
1. Record candidate configuration, runtime version and active generation.
2. State the expected journey before and after membership changes.
3. Identify the authoritative state source and its failure behavior.
4. List representative request durations and remaining capacity.
5. Define the drain deadline and treatment of work exceeding it.
6. Observe new traffic and old requests as separate populations.
7. Reconcile uncertain outcomes by approved correlation identifiers.
8. Assign an owner, next checkpoint and escalation for each exception.
9. Confirm rollback generation without assuming business rollback.
10. Let RUN explain and repeat recovery before accepting handover.
Evidence table:
Population | Expected result | Observation | Gap | Owner | Next checkpoint
New requests |... |... |... |... |...
Existing sessions |... |... |... |... |...
Long requests |... |... |... |... |...
Uncertain outcomes |... |... |... |... |...
Suggested English update:
New traffic is healthy; two outcomes remain under reconciliation
with named owners. The next update will include confirmed status
or the remaining evidence gap for each identifier.
Recorded lab scope: one NGINX worker, synthetic HTTP loopback requests.
Production deadlines, representative load, durable state and actual
business cancellation remain outside the executed fixture.
During a fictional window, a new request reaches B and an old one loses its connection at the deadline. APS checks the ID before authorizing replay and hands outstanding work to the next shift.
Common pitfalls
Choosing a deadline from the teaching example, declaring cancellation from TCP failure or closing the change using only new probes; confusing proxy rollback with reversal of effects.
Related topics: Algorithms, affinity, and state · Retries, limits, and client origin · Draining, releases, and operations
Drain needs criteria for new traffic, old work and uncertain outcomes. Connection deadlines and business outcomes require different evidence.
Reference: NGINX worker shutdown deadline · BigSavant load balancing 2026-09; selected NGINX, HAProxy 3.2, Kubernetes and AWS ALB behavior