Start with the failure the business must withstand
Distinguish component unavailability, regional loss, and logical data error. A replica can help resume after a regional failure while already containing the erroneous deletion that must be undone. Before choosing a mechanism, define the valid state to recover and how much work may be lost. Name owners for failover decisions and reconciliation acceptance. In a funds context, technical resumption can precede validated closing; reporting should make that distinction visible.
Give the client the right listener and consistency
For a failover group, use the listener contract appropriate to the work. A writing application should follow the current primary and handle reconnection; a fixed physical hostname can keep targeting the former destination. An isolated report tolerating lag can use the geo-secondary. Immediate transfer confirmation should not infer failure merely because the replica still shows the old balance. Define which queries must reflect recent writes and which tolerate lag so read optimization does not change business semantics.
Prepare dependencies outside application tables
Data existing on the secondary does not guarantee application login. If users map to master logins, prepare and validate those objects at the destination. Contained users can reduce this dependency but do not remove DNS, networking, or application controls. Include alerts that must watch the new primary. Rehearse with the real operational identity or an equivalent approved test identity. An administrator being able to execute SELECT does not establish service resumption for intended consumers.
Measure RTO and RPO with explicit assumptions
In the model, detection takes four minutes and decision three. Data and networking then start in parallel, taking twelve and eight; final validation takes five. Total time is 4+3+max(12,8)+5=24 minutes. If teams cannot work in parallel, the calculation changes. For RPO, a confirmed point seven minutes before the incident does not establish a five-minute target. A one-hour grace period likewise does not by itself support a twenty-minute RTO. Document what is measured, estimated, or still dependent on authorization.
Restore without losing the ability to reconcile
PITR creates a separate database; plan its validation and the application target. In the 15:20 incident, the 15:19 point can recover deleted rows while excluding valid later movements. Preserve current-state evidence and compare by business identity before deciding replacement or selective repair. Do not repeat external effects merely because they are absent from an old copy. Define who accepts the reconciled balance and how exceptions are documented. Restore is one procedure step rather than an automatic decision about which movements are valid.
Accept resumption, lag, and failback separately
The service can accept requests again while still carrying backlog. If one hundred jobs arrive per minute and one hundred sixty complete, net capacity is sixty; 1800 pending require thirty minutes at constant rates. Measure that lag separately from initial recovery. Confirm capacity and monitoring in the recovery region and avoid rushed failback just because the former region returns. Returning should have suitable synchronization criteria, ownership, and rehearsal. These calculations are planning models and do not establish real Azure subscription timings.
SQL promotion ends in nine minutes, but a login is missing and the alert watches the old region. The rehearsal continues until useful operation and RUN access are demonstrated.
Common pitfalls
Using a replica as historical backup; confusing read-only with immediate consistency; ignoring logins outside the database; excluding detection or validation from timing.
Related topics: RTO, RPO, and capacity · Consistency and reconciliation
Recovery means producing the agreed outcome again with acceptable data, functional access, and observable operation.
Reference: Disaster recovery architecture · AZ-305 objectives 2026-04-17