Concept and mechanism
Load balancing distributes work among existing resources; it does not create capacity lost during failure. If three instances measured at one hundred requests per second serve two hundred and fifty, losing one leaves two hundred and a shortfall of fifty. The example holds workload and limits constant. In production, measure capacity under the actual workload, including dependencies and acceptable latency. Healthy instances can collectively lack sufficient headroom. Equal counts also do not imply equal work: some requests consume more CPU, memory, or time. Observe distribution by target and operation class before concluding that the algorithm or weights represent cost well.
Guided application
Retries can multiply demand. Three total client attempts and up to three proxy attempts for each outer attempt permit nine backend attempts. The prompt includes the first attempt in each total, avoiding confusion between extra retries and attempts. During overload, excluding targets can shift work onto less capacity and start a cascade. Assess admission, priorities, and resource recovery alongside health. Unlimited queues merely accumulate work and missed deadlines. Autoscaling also has time to service: creating pods does not prove readiness, warmup, or available capacity. In a fictional case, mitigation controls nonurgent work while recovering the missing instance.
3 × 100 − 100 = 200 requests/s available; demand of 250 leaves a shortfall of 50.
Common pitfalls
Health as sufficient capacity; weights as resources; layered retries as addition; created pods as ready pods.
Related topics: Routing and TLS boundaries · Algorithms, affinity, and state · Health checks and readiness
Plan capacity after failure and bound amplification that recovery can cause.
Reference: Google SRE handling overload · DR load balancing 2026-09; selected NGINX, HAProxy 3.2, Kubernetes and AWS ALB behavior