← System Design: make decisions, size, and recover systems
05 / 6 · 40 MIN

Failures and recovery

Design for failure domains and demonstrate complete-service recovery.

Concept and mechanism

Redundancy helps only against failures that components do not share. Two instances may depend on the same storage, identity service, network, or faulty configuration. Identify failure modes, required dependencies, and remaining capacity. Observed availability, recovery objectives, and tolerated data loss are different measures. RTO defines tolerated unavailability within the agreed scope; RPO defines the tolerated data-loss window. Having backups or replicas does not demonstrate that these objectives will be met. Recovery must include data, application, access, integrations, and functional validation by someone able to recognize the correct outcome.

Guided application

In an original exercise, RTO is 90 minutes and RPO is 15. Recovery takes 70 minutes, but the only recoverable point is 03:00 for a failure at 03:22. Assuming no logs or other later recovery exist, elapsed recovery meets RTO while the 22-minute loss window fails RPO. Do not mark the exercise wholly successful. Record gaps and owners, confirm destination capacity, and prepare return after operating at the alternate location. In APS scenarios, include batch dependencies and in-flight files: recovering servers without reconciling work can leave business state inconsistent.

IN PRACTICE

70≤90 meets the deadline; 22>15 exceeds tolerated loss in this exercise.

Common pitfalls

Two replicas as no common failure; backup as proven restoration; met RTO as met RPO.

Related topics: Requirements and architecture decisions · Capacity and latency · Data and consistency

Take this idea with you

Validate recovery against each objective and final business state.

Create account

Reference: Defining reliability targets · System design patterns; PostgreSQL18 scoped examples; primary guidance consulted 2026-09-30