Draw states before timers
A fictional positions application requests more instances before close. The operating system starts, but the service is still fetching configuration, filling caches, and opening connections. Counting instances does not establish that all can process requests. Draw a timeline covering launch, bootstrap, load-balancer registration, health, stabilization, and exit. For each transition, define who observes it and what evidence is required. A health check returning 200 without checking minimum configuration can admit unusable capacity. The project team should agree the readiness contract with APS and development, including behavior when dependencies are unavailable and ownership of exception acceptance during a change window.
Separate grace period, warmup, and hook
Grace period avoids premature health-check replacement while a new instance initializes in InService; it does not itself block ALB traffic. It also does not prevent replacement if the instance leaves EC2 running state. Default instance warmup addresses stabilization of aggregated metrics used by dynamic scaling. It does not replace a bootstrap gate. A launch lifecycle hook can hold the instance before admission. Use the three according to the observed problem rather than increasing every timer in search of stability. Measure how long configuration takes and how long load takes to settle, because those intervals can differ and have different causes.
Give the hook an honest outcome
If bootstrap fails to obtain mandatory configuration, sending CONTINUE turns a known failure into apparently available capacity. Retain diagnostics and use the appropriate outcome, correcting the cause before repeating launches without limits. Heartbeats maintain waiting within deadlines but do not establish functional progress, and a global deadline exists. The runbook needs escalation and a terminal state. Under default termination behavior without additional retention policy, ABANDON does not cancel deletion: it allows termination and prevents remaining actions. Additional retention mechanisms require actual configuration to be checked. Do not carry a launch-action interpretation into a termination action.
Connect health signals to the right decision
Attaching a target group does not mean Auto Scaling uses ELB results to replace instances. Enable that integration when it belongs in the design and confirm endpoint meaning. A 401 response after an authentication change should not be hidden by broadening the matcher to every code. Investigate access and the health-check contract. Also consider common failure: if every registered target is unhealthy across all enabled AZs, ALB can route in fail-open mode. Replacing copies dependent on the same unavailable service may then merely consume capacity. Connect the shared cause to containment and recovery of the business path.
Plan exit and load distribution
During deregistration, draining does not prove every request finished. The process must respect in-flight work and deadlines, with recovery for incomplete effects. The state can also remain visible until timeout even with no active requests; inspect concrete evidence. Scale-in protection selects instances eligible for reduction but does not protect them from every failure. If all are protected, desired capacity can decrease while the observed fleet remains larger. For gradual admission, check attribute combinations: slow start is unsupported with least outstanding requests or weighted random. Algorithm choice, readiness, and warmup should be exercised together.
Measure headroom after failure
The local model sums validated capacity by AZ and removes the lost zone. With two 100-request/s instances in each of three zones, losing one leaves 400 requests/s. A load of 450 exceeds that capacity even if desired capacity still indicates six instances. The exercise neither calls AWS nor simulates automatic recovery. Use it to discuss headroom before additional capacity exists, quotas, startup time, and dependency limits. In a real rehearsal, also measure latency, queues, and errors with one zone unavailable. The RUN conclusion should combine surviving capacity, admission and exit behavior, and action criteria when demand exceeds headroom.
# Original local model: assumes validated independent capacity per zone.
def surviving_capacity(zones, lost):
if lost not in zones:
raise ValueError("Unknown failure zone")
if any(type(v) is not int or v < 0 for v in zones.values):
raise ValueError("Capacity must be a nonnegative integer")
return sum(v for zone, v in zones.items if zone!= lost)
zones = {"zone-a": 200, "zone-b": 200, "zone-c": 200}
assert surviving_capacity(zones, "zone-b") == 400
assert surviving_capacity(zones, "zone-b") < 450
assert surviving_capacity({"only-zone": 200}, "only-zone") == 0
try:
surviving_capacity(zones, "unknown")
raise AssertionError("Unknown zone accepted")
except ValueError:
pass
Expansion requests four instances, but all fail the same configuration dependency. Correcting readiness and bootstrap recovers useful capacity; counting InService does not resolve failure.
Common pitfalls
Confusing grace period with a traffic gate; using warmup as bootstrap; interpreting ABANDON as termination cancellation; relying only on instance count.
Related topics: Recovery rehearsals and functional evidence
Usable capacity requires configuration, health, stabilization, and headroom. Each control covers part of that path.
Reference: Auto Scaling health check grace period · DOP-C02