Concept and mechanism
Resilience depends on the complete service path. Workers in two AZs remain vulnerable if they need a single server in the first. Identify configuration, data, naming, and authorization dependencies and rehearse the relevant failure domain. Distinguish a running process from a healthy service: an ECS task in RUNNING can fail load-balancer health checks. For queue-worker sizing, backlog per instance relates visible messages to InService instances. An initial target can come from acceptable waiting time divided by average processing time. The calculation is a model, not a promise about every message or immediately available capacity.
Guided application
With 60 seconds of acceptable waiting and 0.5 seconds average processing per message, the target is 120. A queue of 960 messages with four workers has 240 per worker. Before expanding, confirm limits, dependencies, and actual effects. Protect long tasks during scale-in and release that protection when appropriate; do not confuse it with protection against every failure. During recovery, add sequential steps until the service is validated: 18 minutes restoring, seven configuring, and ten validating total 35. AWS Backup restore testing can support rehearsals and functional validation, but job completion does not prove business success. Preserve tags required to clean up test resources and monitor cleanup completion.
The committee requires a 30-minute RTO. A complete 35-minute rehearsal exposes a gap even if restoration alone takes 18.
Common pitfalls
Average as SLA; two AZs as absence of single dependencies; restoration as complete recovery; tests without cleanup.
Related topics: Telemetry and useful alarms · Events, incidents, and replay
Measure work per capacity and recovery through to functional outcome.
Reference: DOP-C02 domain 3 · DOP-C02