1. Separate flow, session and transaction
An application can retain a logical session for hours while network components retain state according to their own tuples and timeouts. GWLB associates flows with appliances using five-tuple by default. Two- or three-tuple modes are incompatible with TGW appliance mode. UDP also has state in the load balancer, but its 120-second timeout is not configurable. A message after prolonged silence can be treated as a new flow. For TCP, increasing GWLB idle timeout requires checking appliance ENI tracking; a lower ENI limit can remove state first. At design review, request real examples of idle periods rather than relying only on continuous probes that never exercise expiry.
2. Choose recovery with explicit limits
By default, no_rebalance keeps existing flows on the previous target when it fails or is removed, while new flows use healthy targets when available. Rebalance changes destination but does not automatically replicate appliance state. Failover attributes for unhealthy and deregistration must match. If the strategy depends on explicit TCP reconnection, assess reset options with five-tuple and no_rebalance; do not combine them with rebalance or promise an equivalent UDP effect. Health-check detection time is only part of recovery time. Propagation, retransmissions, retries, and reconciliation can extend impact. Define acceptance for old and new sessions and confirm the business outcome after the induced failure.
3. Validate path and packet size
GWLB asymmetric-flow support has limits: the load balancer needs to observe the initial packet in the documented case; a path bypassing it outbound and passing through it only on return is unsupported. Maintain required inspection in both directions with identified tables and endpoints. For MTU, the appliance segment includes 68 GENEVE bytes beyond the original packet. An 8500-byte packet requires at least 8568 on that segment. GWLB does not fragment IP or provide PMTUD through the fragmentation-needed ICMP message. Test realistic sizes and identify the segment limiting the flow. If the project needs to remove a subnet from an existing GWLB, include a new load balancer and migration in the plan.
4. Make health checks meaningful
Health checks are probes with their own configuration and path. GWLB checks are distributed and use consensus; several requests in an interval do not prove an attack. On NLB, an HTTP probe uses Host with node IP and listener port, potentially selecting a different virtual host from a normal application request. If the endpoint supports only TLS 1.3, check the documented HTTPS-health-check limitation before blaming the application. The NLB SG controls outbound health checks, while the target SG must allow both service and probe. A static page returning 200 remains limited evidence. Combine technical health with a representative synthetic transaction while making its cost and production effects explicit.
5. Treat source metadata as a contract
Proxy Protocol v2 adds binary information to a connection. Before enabling it, confirm that the backend and health endpoint can parse the header; it is also sent on probes without actual client-connection information. An HTTP 400 after the change can indicate an incompatible parser. Broadening the matcher to accept that error does not fix the contract. With a TCP listener, prior headers may exist, and an alternate path to the backend can present a header that did not come from NLB. Constrain trusted proxies. If a target instance calls its own internal NLB with preserved IP, consider the hairpin limitation. Changing preservation may require Proxy Protocol to retain observability, creating a dependency that must be rehearsed.
6. Distinguish distribution, isolation and rollback
An NLB created without an SG cannot receive its first SG later; anticipate this migration dependency. Where the NLB has an SG, you can reference it in target SGs even with client-IP preservation. For PrivateLink subject to inbound rules, the source considered is the client private IP, not the endpoint-interface IP. Health checks are not isolation: when every target fails, NLB can fail open and GWLB can still select an unhealthy appliance. Define explicit controls to contain a compromised service. On ALB, weights between target groups do not provide automatic failover from an empty or unhealthy group to another healthy one. A canary release needs rollback signals and a routing change that has actually been rehearsed.
# Local compatibility illustration, not an AWS API validator or packet capture.
def tcp_reset_compatible(stickiness_enabled, failover):
return stickiness_enabled is False and failover == "no_rebalance"
def encapsulated_size(original_bytes):
if not isinstance(original_bytes, int) or not 0 < original_bytes <= 8500:
raise ValueError("exercise expects an original packet between 1 and 8500 bytes")
return original_bytes + 68
assert tcp_reset_compatible(False, "no_rebalance")
assert not tcp_reset_compatible(True, "no_rebalance")
assert not tcp_reset_compatible(False, "rebalance")
assert encapsulated_size(1500) == 1568
assert encapsulated_size(8500) == 8568
print("five compatibility and size cases passed; no AWS changes made")
GWLB retains state for 900 seconds but the ENI for 350. A session idle for 500 seconds can fail despite the larger load-balancer timeout.
Common pitfalls
Rebalance treated as state replication; health checks as isolation; a broadened matcher as a fix; ALB weights as automatic failover.
Related topics: Failover and middleware sessions · Health checks and observability
A change is accepted only when routing, protocol, session and transaction recover within agreed criteria.
Reference: Gateway Load Balancers · ANS-C01