← Load balancing: traffic and resilience
07 / 8 · 60 MIN

Attempts and passive health with NGINX

Reconstruct a request path and distinguish recovered success, temporary exclusion and replay risk.

Prepare a controlled comparison

The lab uses locally built NGINX 1.30.5, one worker and two fictional HTTP servers, A and B, on loopback. Backends record method, path and received body. Each failure test has its own upstream group so it does not inherit exclusion caused by the previous test. This separation matters: if A is already excluded, a POST can go directly to B and no longer test the intended guard. Before interpreting a result, identify configuration, request sequence and each group’s initial state.

Read both perspectives of the same request

On the first failure GET, A returns 503 and B returns 200. The client receives 200 while the proxy log retains upstream_status of 503, 200 and both addresses in matching order. The query therefore remained available through recovery, but A did not become healthy because the final status was positive. For a fictional positions portal, also collect latency and demand by target. A global success indicator can remain good while each recovery adds work and reduces headroom on the responding backend.

Observe exclusion without inventing recovery

The group configures A with max_fails=1 and fail_timeout=60s, B as backup and http_503 as a retry condition. The next GET reaches only B. The runner observes this immediate consequence but does not wait for A to reenter. Recording “automatic recovery validated” would exceed the evidence. The next operational question is when A is selected again and whether it can perform the functional query. Passive state also represents neither an application restart, a periodic check nor proof that all dependencies are available again.

Compare guarded POST and explicit replay

Two independent groups receive the same synthetic body. Without non_idempotent, A receives the POST and returns 503; B does not receive it. In the other group the option allows replay, and events show the body at A and B, ending with 200. Servers only record receptions and respond; they implement no financial transactions. Therefore two HTTP receptions were observed, not two transactions or exactly-once execution. For a real API, review identifiers, persisted effects and reconciliation before turning this result into a retry policy.

Count attempts and bound distribution claims

In a third failure group, proxy_next_upstream_tries 1 produces only an A observation, with final response 503. In this experiment, one means one total attempt, not one additional attempt. A separate test sends eight short GETs with weights three for A and one for B, obtaining six A responses and two B responses. This confirms distribution in that controlled sample. It measures neither maximum capacity, concurrency, long sessions nor behavior when a target becomes ineligible. Keep share calculations separate from any production capacity commitment.

Apply the diagnosis in an incident meeting

In a fictional incident, the manager sees final 200s and asks whether it can be closed. Explain the A503, B200 sequence, observed exclusion and missing recovery evidence. Propose assessing remaining capacity and containing the failing instance without causing a cascade. Provide ordered events, configuration and experiment limits. The summary is to distinguish recovered availability from each target’s health, HTTP reception from business effects, and distribution from capacity. Use the questions to decide which evidence to collect before the next change.

# Synthetic local observations in content/labs/lb-upstream/evidence.json
# retryThenExclusion: A -> B; next request B only
# attemptLogs: upstreamStatus "503, 200" then "200"
# postGuard: A only; explicitReplay: A and B
# attemptLimit: tries=1, one observed backend attempt
IN PRACTICE

Fictional case: queries finish at B after 503 at A. The team keeps the incident open, confirms remaining capacity and investigates A using the attempt sequence.

Common pitfalls

Confusing final 200 with every target being healthy, max_fails with restart, reception with effect, and eight sequential requests with a load test.

Related topics: Health checks and readiness · Retries, limits, and client origin · Draining, releases, and operations

Take this idea with you

Reconstruct statuses and targets per attempt before concluding health, recovery or replay safety. The lab bounds exactly what was observed.

Create account

Reference: NGINX HTTP proxy module · BigSavant load balancing 2026-09; selected NGINX, HAProxy 3.2, Kubernetes and AWS ALB behavior