Define what recovery means
A fictional platform receives instructions, stores documents in S3, and updates a database. During recovery rehearsal, the team can read older documents, but no new instruction is processed. The result shows why restoring data is only part of acceptance. Define critical operations, tolerable data loss, recovery deadline, and decision authority before the window. Include application versions, configuration, permissions, and integrations in the dependency inventory. Record who validates each result and when the service may accept traffic. A read test can be necessary without being sufficient: writing, event emission, and reconciliation of the interrupted interval may also be required. These checks should follow the service contract.
Choose each RDS copy’s role
Distinguish a read replica from a standby in a Multi-AZ DB instance deployment. The former receives asynchronous replication and can serve reads; the latter supports availability and is not a read-query target in that architecture. Do not generalize this restriction to Multi-AZ DB clusters, which have a different design. An application requiring immediate reading of an acknowledged write may receive older state from a lagging replica. Define which queries tolerate that difference. During Multi-AZ failover, the endpoint is updated through DNS and existing connections need recovery. A JVM retaining the previous address may fail while new clients connect, so the runbook should cover DNS caching, pools, and reconnection.
Prepare a planned promotion
In the RDS PostgreSQL example, the source remains reachable and no acknowledged writes may be lost. Before promotion, control new writes and confirm the replica received pending changes. Promotion stops replication to that replica and creates a standalone instance; it does not retain two synchronized writers. Plan routing, permissions, backups, and the later topology. Validate functional operations before declaring the change complete. If the source is lost and lag is unknown, the decision has different uncertainties: do not reuse the promise from this planned rehearsal. The change owner should record evidence, remaining risk, and the condition preventing continuation. Schedule pressure does not change replication properties or justify an unsupported guarantee.
Prove S3 destination state
A new live replication rule does not automatically resolve every earlier object. Use Batch Replication where appropriate, with a manifest and per-object, per-version report. A Complete job can contain failed tasks; compare outcomes with the expected set. With multiple destinations, one observed copy does not prove success everywhere. If replication failed because of KMS permissions, fixing permission is insufficient to assume automatic retry of FAILED objects. Confirm the destination key too: acceptance of PutBucketReplication does not validate the supplied key. Finally, bucket configuration such as notifications does not accompany object replication. Explicit deletion of a source version is not propagated in that way either, requiring a separate retention and cleanup policy.
Restore DynamoDB and dependencies
PITR restores into a new table. Source operations may continue, so define the intended instant and reconcile later changes before cutover. Query the actually recoverable interval; increasing retention today does not retroactively create days that were not retained. ContinuousBackupsStatus and PointInTimeRecoveryStatus have different meanings and must not be confused. After restore, check settings requiring reconfiguration, including Streams, TTL, autoscaling, alarms, tags, and PITR. If an index is excluded to reduce recovery work, test queries that depended on it. An Active table may be available without satisfying the application contract. Use the new table and stream identities when validating permissions and consumers, avoiding stale references in deployment configuration and operational instructions.
Exercise: control cutover order
The local model represents a planned change through states: authorization, write control, synchronization, promotion, dependencies, and functional validation. It rejects out-of-order steps and accepts traffic only after completing the sequence. It does not query RDS, measure lag, guarantee isolation, or implement real failover. Run the cases, try promoting before synchronization, and describe real evidence supporting each step. Then add a stop condition if functional validation fails. In the rehearsal report, record observed timestamps and compare them with the internal objective without turning one exercise’s duration into a vendor guarantee. Shift handover should state the final topology, outstanding work, and reconciliation owner, with enough detail for the next operator to act.
# Original local ordering exercise, not an RDS failover engine or atomic cutover guarantee.
STEPS = ("authorized", "writes_controlled", "caught_up", "promoted", "dependencies_ready", "function_validated")
def advance(done, step, evidence):
if tuple(done)!= STEPS[:len(done)] or len(done) >= len(STEPS):
raise ValueError("invalid or finished sequence")
if step!= STEPS[len(done)] or evidence is not True:
raise ValueError("order or evidence missing")
return done + [step]
def may_accept_traffic(done):
return tuple(done) == STEPS
done = []
for step in STEPS:
assert not may_accept_traffic(done)
done = advance(done, step, True)
assert may_accept_traffic(done)
for previous, step, evidence in [([], "promoted", True), (["authorized"], "writes_controlled", False), (["promoted"], "dependencies_ready", True)]:
try:
advance(previous, step, evidence)
except ValueError:
pass
else:
raise AssertionError("unsafe sequence accepted")
Fictional example: documents exist at the destination, but notifications do not. The team restores integration and reconciles files received during interruption before accepting the service.
Common pitfalls
Pitfalls: confusing standby and read replica, promoting with unreconciled writes, accepting Complete without per-object outcomes, and assuming restore reinstates every integration.
Related topics: Security coverage and compliance evidence
Accepted recovery requires data suited to the objective, restored dependencies, and demonstrated functional operation at the destination.
Reference: DOP-C02 resilient cloud solutions objectives · DOP-C02