← AZ-305: Azure architecture and production decisions
05 / 8 · 30 MIN

Continuity and demonstrated recovery

Design and measure operational recovery, including dependencies.

Concept and mechanism

RTO limits time until the defined operation is restored; RPO limits the time gap in recoverable state. Starting infrastructure proves neither target. The recovery path includes networking, identities, data, application, and validation while respecting dependencies. Sequential stages add duration; parallelism is valid only when the design permits it. Zonal distribution is not equivalent to complete-region recovery. Asynchronous replication should not be presented as a universal zero-loss guarantee either.

Guided application

Drills should demonstrate recovery-asset access, destination capacity, and functional results. Green backups need validated restoration. A topology can have active frontends in two regions while still depending on one writable database. Also document failback: data and operations produced in the secondary need handling before returning. Record gaps and owners without unilaterally changing business commitments.

IN PRACTICE

Network 8 min, database 22 min, application 8 min, and validation 12 min total 50 min. A45-minute RTO is exceeded by 5 minutes.

Common pitfalls

Excluding validation from timing; comparing RPO with restore duration; keeping runbooks only in the unavailable region.

Related topics: Compute and operational responsibility · Integration and change contracts

Take this idea with you

Recovery is a rehearsed capability of the complete service.

Create account

Reference: Disaster recovery architecture · AZ-305 objectives 2026-04-17