← AWS Solutions Architect Associate: architecture decisions
04 / 10 · 25 MIN

Resilience, capacity, and recovery

Relate possible failures to remaining capacity and business objectives.

Concept and mechanism

Distributing instances across Availability Zones reduces dependence on one location, but remaining capacity and shared dependencies still constrain recovery. A load balancer can only route to healthy, available targets. Auto Scaling helps adjust capacity, but startup time, quotas, availability, and application state affect the response. Sessions stored only in one instance’s memory may be lost when it leaves service. Architecture should address how required state survives replacement and how the application handles repeated requests.

Guided application

RPO expresses tolerated data loss in time; RTO expresses target recovery time. Backup and restore, pilot light, and warm standby have different costs and operational steps. Compare them with the agreed need and measure a complete exercise, including credentials, network, data, DNS, and functional validation. DNS failover changes future responses according to configuration; caches and existing connections influence client-observed timing. The runbook needs activation criteria, owners, communication, and return-to-normal steps.

IN PRACTICE

Two AZs have four workers each and peak demand requires six. Losing one leaves four: a two-AZ diagram does not establish adequate capacity during failure.

Common pitfalls

Measuring infrastructure alone; confusing backup frequency with RTO; ignoring post-failure capacity; promising instant DNS change.

Related topics: Storage and content delivery · Databases, replicas, and cache

Take this idea with you

Resilience needs observed capacity and recovery as well as redundant components.

Create account

Reference: Disaster recovery options · SAA-C03