Concept and mechanism
Retrying another backend can recover a transient failure but can also duplicate an already-applied operation. For a POST sent before timeout, establish the contract and uncertain outcome. In NGINX, allowing non_idempotent permits behavior that does not create application deduplication. Define attempt and time limits coherent with the client. If it gives up after 50 seconds while the proxy continues to 60, work can remain without a waiting consumer and with an unknown outcome. Cancellation and operation effects need their own analysis. The largest timeout on one layer does not replace limits on the others.
Guided application
Also distinguish limit scopes. NGINX upstream keepalive 32 concerns an idle-connection cache per worker rather than every connection it may open. Without a shared memory zone, max_conns applies per worker: four limits of ten can permit up to forty active connections in the example. Idle connections introduce further nuances. For client origin, X-Forwarded-For is not trusted merely because it resembles an IP. Define authorized proxies and parsing rules while controlling direct paths bypassing that boundary. In a fictional case, the same field is accepted through the proxy and directly; fixing only the parser does not resolve who may supply the claim.
Four workers with a local limit of ten do not equal one global counter of ten.
Common pitfalls
A new backend as deduplication; the largest timeout as a global deadline; idle cache as a total limit; a header as authenticated identity.
Related topics: Routing and TLS boundaries · Algorithms, affinity, and state · Health checks and readiness
Establish the scope and consequence of each limit and origin claim.
Reference: NGINX retries, timeouts and upstream forwarding · DR load balancing 2026-09; selected NGINX, HAProxy 3.2, Kubernetes and AWS ALB behavior