← Disaster Recovery: prepare, recover, and validate
08 / 8 · 60 MIN

Resumption: replay and operational acceptance

Reconcile effects that survived the disaster and decide when the recovered service can enter operation.

External state does not move backward with the backup

The fictional producer stores an order as pending before backup. Later, the synthetic receiver accepts the order and the producer marks it sent. Restoring the old copy makes the producer show pending again, but the receiver still has a confirmed effect. In the lab, both are separate local databases; there is no call to a real partner. The example shows why a restored marker is not global truth about delivery. In batch recovery, identify states outside the restored set, preserve identifiers, and consult authorized confirmations. Automatically restarting every pending item can repeat work that has already produced an effect.

A retry needs identity and content

The model receiver binds order-1 to amount 100. A retry with the same identifier and content returns duplicate while retaining one effect. The same identifier with 101 is rejected as a conflict. A new identifier with 100 produces another effect: the receiver cannot know that the previous operation was intended to be retried. These rules are an explicit model contract, not a universal guarantee of banking APIs. They also assume history still exists. If backup recovery goes beyond deduplication retention, the key alone may be insufficient. Define the replay window, available evidence, and conflict handling before authorizing resubmission.

Write during an exercise without a real effect

The write probe inserts a third order at the isolated destination, observes it inside the transaction, and rolls back. Two orders then remain and reconciliation passes. The lab also confirms that the source and backup artifact stayed unchanged. This demonstrates a reversible local operation with the test process privileges. It does not validate IAM, keys, partner connections, or the actual RUN team account. Furthermore, SQL rollback does not automatically undo an external request already sent. In an application exercise, control output destinations before enabling consumers and use operations whose effects and cleanup are defined in the approved plan.

Measure through the agreed function

In a fictional exercise, unavailability starts at 09:00. Restore ends at 09:20, configuration at 09:28, and functional validation at 09:35. If the commitment is validated service within 30 minutes, the result is 35, with a five-minute overrun. Retain intermediate milestones to understand the path delaying recovery. Do not select the most favorable milestone after the exercise. During planning, consider overlap: a 20-minute restore and 12-minute identity preparation can run in parallel; eight-minute validation dependent on both finishes at minute 28, absent other waits. That calculation is an estimate and needs observation in the relevant environment.

Accept capacity and operation as well as data

A correct database does not guarantee capacity for incoming work. In the fictional example, 50 orders arrive each minute and the recovered destination processes 20. The queue grows by 30 per minute even if every processed order is correct. The team should report that degraded mode, estimate cut-off impact, and discuss containment or additional capacity. The exercise also needs the identity that will operate the service: a privileged personal account can hide missing permissions in the RUN account. Tie each conclusion to what was observed. The two-order lab does not establish production throughput; it only provides a structure for more complete acceptance questions.

Close with protection and residual risk visible

The recovery report should distinguish restored function, completed reconciliation, objectives met or exceeded, and protection still pending. If the new destination accepts writes but lacks a validated backup, the first recovery has not resolved exposure to the next failure. Record owner, deadline, expected evidence, and who may accept temporary risk. An approved degraded mode does not mean pretending risk is absent. In the fictional handover, the team delivers reconciled identifiers, check results, timing milestones, and outstanding actions. After correcting a gap, repeat the corresponding check; the existence of a closed task does not replace observation of the result.

original ID + original content -> duplicate, one effect
original ID + changed content -> conflict
new ID + original content -> second effect
# Synthetic local receiver contract, not a universal API guarantee.
IN PRACTICE

The restored producer shows pending for an already executed order. Retrying the original ID retains one effect; a new ID creates a second effect at the synthetic receiver.

Common pitfalls

Resending with new IDs; assuming local rollback undoes external effects; measuring only restore; treating a privileged account as proof of RUN autonomy.

Related topics: Restore: recovered point and integrity · RTO, RPO, and dependencies · Exercises and functional validation

Take this idea with you

Resumption requires reconciling what happened outside the backup and explicitly accepting function, time, capacity, and protection.

Create account

Reference: Test disaster recovery implementation · BigSavant recovery 2026-09; PostgreSQL 18, etcd 3.6 and selected AWS/Azure behavior