← AWS Solutions Architect Associate: architecture decisions
10 / 18 · 55 MIN

Recovery: dependencies, capacity, and cost

Measure the recovery critical path and compare designs under the same service commitments.

Define the service that must return

For a fictional positions service, recovery means more than starting a database. Define accepted reads and writes, which reconciliation confirms balances, and who authorizes reopening. RTO needs an agreed measurement boundary. If it starts at failure, detection and decision belong in elapsed time. RPO describes the tolerated data point and must be compared with recoverable evidence. A continuously configured replica can lag. Also record in-flight messages and externally initiated operations: one database timestamp may not explain the complete state of the business process.

Find dependencies crossing Regions

Walk through recovery-service configuration and identify references to the primary Region. A replicated secret does not help if the application still calls the original endpoint. Validate the regional copy and access to its protecting key. With KMS multi-Region keys, policies and grants need regional treatment; related material does not mean synchronized authorization. Add images, configuration files, and external integrations to the dependency list. For every entry, ask whether the destination starts when the origin is unavailable. The answer should be supported by an appropriate exercise without relying on broader credentials than those intended for production.

Calculate the sequence with explicit parallelism

In the teaching exercise, data and networking start together. Data takes eighteen minutes and networking seven. Configuration starts only when both finish and takes four; subsequent validation takes six. This plan takes max(18,7)+4+6, or twenty-eight minutes. If three earlier minutes of detection and decision belong in RTO, total duration is thirty-one. Neither add parallel tasks nor remove mandatory steps to improve the number. Measurement identifies where an improvement actually changes the outcome: reducing networking from seven to four does not shorten this critical path.

Distinguish a working service from sufficient capacity

A warm standby can be functional at reduced capacity. In the example, it sustains one hundred and twenty requests per second, but failover requires four hundred and scaling takes twelve minutes. The team needs advance capacity or agreed admission control and controlled degradation during that window. Do not conclude that every user can enter because a simple transaction passed. If the design creates components during the incident, identify control-plane dependencies, quotas, and resource availability. Pilot light and warm standby are choices to establish against the service objective, not labels that automatically guarantee RTO.

Protect consistency before reopening writes

A DNS change is not, by itself, database promotion or exclusion of an old writer. Before reopening writes, establish which destination has authority and how the other is prevented from producing incompatible changes. The specific mechanism depends on the technology and needs exercising. If the failure is logical corruption already replicated, promoting the copy can preserve the problem. Recovery then needs a clean point, reconciliation, and handling of later legitimate operations. Separate continued access from correct state. Record the business decision, because quickly restoring access to an incorrect balance is not accepted recovery.

Compare costs under equivalent commitments

FINOPS proposes shutting down components that spend little time busy. Before accepting, measure the resulting recovery plan against the same RTO, capacity, and data requirements. In a separate fictional model, A=10+0.05V and B=30+0.01V cost the same at five hundred GB. Below that volume A is cheaper; above it B is cheaper. The conclusion depends on stated costs and equivalent requirements. If a shared Resolver remains necessary, retiring one consumer does not eliminate its entire bill. Close the analysis with a volume assumption, genuinely removable resources, and a usage review date. The summary combines complete duration, correct data, sufficient capacity, and demonstrable marginal cost.

IN PRACTICE

Guided case: service returns in 31 minutes against a 30-minute RTO. Reducing networking does not remove the delay because data dominates the parallel phase. Propose a critical-path improvement without removing functional validation.

Common pitfalls

Omitting detection from RTO; promoting corrupted data; confusing DNS with fencing; cutting fixed cost without measuring the new startup; treating KMS grants as global.

Related topics: RTO, RPO, and critical path · KMS and regional secrets · Capacity, FINOPS, and reconciliation

Take this idea with you

Recovery is established when the service returns with acceptable data and capacity within the agreed window using available dependencies.

Create account

Reference: Disaster recovery options in the cloud · SAA-C03

AWS is a trademark of Amazon.com, Inc. or its affiliates. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by AWS. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.