← Load balancing: traffic and resilience
10 / 12 · 60 MIN

Passive failure, reentry and readiness

Observe exclusion and reentry, distinguish retry from accounted failure and validate whether a check represents the intended application.

Follow a failure into the next request

The passive group uses A as primary, B as backup, max_fails=1 and fail_timeout=1s. A initially returns 503 and configuration permits retry under that condition. The first request traverses A and B, finishing with 200; the next goes directly to B. Positive final outcome does not erase the first attempt or temporary exclusion. Correlate backend events with upstream_addr and upstream_status in the access log. Upstream groups are separate so that one exercise’s failure does not contaminate another’s state.

Waiting does not verify recovery

The script changes the fictional backend to return success and waits 2.2 seconds without sending requests. No new events appear during that interval. Waiting exceeds configured fail_timeout but performs no active health check. Only the next request demonstrates A’s reentry and a 200 response. In an actual incident, period expiry does not prove the cause disappeared. If the application remains broken, the new attempt can fail again. Recovery needs evidence of the selected instance and functional contract rather than only an expired timer.

Test accounting exceptions

Two groups show why a universal exclusion rule should not be memorized. In noaccount, max_fails=0 disables accounting: two requests traverse A with 503 and B with 200. Retries remain active. In single, only A exists; despite max_fails=1 and fail_timeout=60s, two requests reach A and receive 503, matching the documented single-server exception. Read member count, options and edition before diagnosing actual configuration. Zero, down, lack of alternatives and retry are different concepts.

Separate retry and accounted failure

The notfound group permits retry on http_404. A returns 404 and B returns 200 for two consecutive requests. A is tried again because 404 does not count as an unsuccessful attempt for passive exclusion, although it permits moving to the next server when configured. Do not treat retry conditions as identical to accounted events. Behavior is specific to NGINX and executed options. For APS analysis, document condition, attempts, final outcome and subsequent eligibility while preserving both evidence levels.

Give the check the correct contract

The health backend is a teaching fixture: it returns 200 for Host default.fund.test and 503 for Host app.fund.test. The two proxy routes differ in that header. The experiment shows that a generic response can hide failure of the intended service without exercising the active-check module. For an actual portal, define Host, path, status and any expected content according to the readiness contract. Keep the check inexpensive and interpret shared dependencies. Repeating a check against the wrong virtual host merely increases the number of irrelevant samples.

Hand evidence and remaining work to RUN

Use the worksheet below to simulate an English RUN handover meeting. One person presents per-request outcomes; another checks whether success came from the recovered primary or the backup. Add concurrency limits, worker scope, readiness criteria and actions when every target becomes unavailable. The lab does not exercise cloud, Kubernetes, HAProxy or production load. Those implementations need their own evidence. Record owners and closure criteria for gaps. The workshop is proposed for study; no human workshop or independent specialist review has taken place.

PROPOSED LOAD-BALANCING HANDOVER
Scope: NGINX version/build, workers, shared state, keepalive configuration.
Selection: method, weights, eligible targets, active-connection evidence.
Admission: max_conns scope, exhaustion response, remaining headroom.
Passive health: counted conditions, max_fails, fail_timeout, exceptions.
Readiness: Host, path, status/content contract and check cost.
Request evidence: final status, ordered upstream attempts, backend events.
Recovery: selected instance, functional result, repeat-failure handling.
Change: stop criteria, recovery action, owner and communication.
Outstanding: representative workload, multi-worker limits and real platform.
Statement: Bounded concurrency and passive-state checks passed;
representative capacity still needs validation.
No human workshop has been performed.
IN PRACTICE

Exercise: a report shows final 200s while A repeatedly fails. Reconstruct attempts, identify max_fails=0 and write an incident update retaining client success and primary failure.

Common pitfalls

Treating waiting as repair, max_fails=0 as no retries, 404 as an accounted failure, or a generic check as application evidence.

Related topics: Algorithms, affinity, and state · Health checks and readiness · Capacity and cascading failures

Take this idea with you

Eligibility follows concrete rules. Operational closure requires identifying the target, observing recovery and validating the contract the user actually uses.

Create account

Reference: NGINX retries and upstream headers · BigSavant load balancing 2026-09; selected NGINX, HAProxy 3.2, Kubernetes and AWS ALB behavior