Concept and mechanism
Distributing instances across Availability Zones reduces dependence on one location, but remaining capacity and shared dependencies still constrain recovery. A load balancer can only route to healthy, available targets. Auto Scaling helps adjust capacity, but startup time, quotas, availability, and application state affect the response. Sessions stored only in one instance’s memory may be lost when it leaves service. Architecture should address how required state survives replacement and how the application handles repeated requests.
Guided application
RPO expresses tolerated data loss in time; RTO expresses target recovery time. Backup and restore, pilot light, and warm standby have different costs and operational steps. Compare them with the agreed need and measure a complete exercise, including credentials, network, data, DNS, and functional validation. DNS failover changes future responses according to configuration; caches and existing connections influence client-observed timing. The runbook needs activation criteria, owners, communication, and return-to-normal steps.
Two AZs have four workers each and peak demand requires six. Losing one leaves four: a two-AZ diagram does not establish adequate capacity during failure.
Common pitfalls
Measuring infrastructure alone; confusing backup frequency with RTO; ignoring post-failure capacity; promising instant DNS change.
Related topics: Storage and content delivery · Databases, replicas, and cache
Resilience needs observed capacity and recovery as well as redundant components.
Reference: Disaster recovery options · SAA-C03