← AWS DevOps Engineer Professional: operations and delivery
14 / 24 · 100 MIN

Recovery rehearsals and functional evidence

Turn a technical restore into a rehearsal measuring data loss, dependencies, permissions, and service recovery.

Define what must return

A valid backup is an input to recovery. The rehearsal’s purpose is to demonstrate that service returns to accepted outcomes within agreed limits. In a fictional payments example, a restored volume can contain intact data while DNS, secrets, and routes still point to unavailable resources. Inventory these dependencies and identify which are rebuilt, restored, or changed. Define who accepts data, performance, and degraded operation. A resource appearing in the console is not the business success criterion. Keep a timeline from failure through validation and distinguish completed technical steps from dependencies still outstanding.

Measure RPO at the usable point

If the source fails at 12:17 and the latest consistent point available at the destination represents 12:04, exposure is thirteen minutes. A 12:15 backup not yet usable at that destination does not automatically reduce that exposure. Compare the result with agreed RPO and analyze the data interval requiring reconciliation. Replication existence is not universal proof either: what was actually replicated and whether the point is consistent for the required set matter. Use fictional or appropriately controlled data in the exercise. Reporting should identify the chosen point, rationale, and read evidence in addition to the job name.

Isolate effects and prepare authorization

A restored application can start jobs and contact real destinations using old configuration. Before starting it, review network, credentials, schedulers, and destinations. A test account helps separate resources but does not establish isolation of every integration. Define validation outcomes that do not submit real payments. In AWS Backup, separate operator authorization to pass a role from permissions of the role used by the service to restore. The role needs actual resource actions, including applicable configuration. Request acceptance does not prove every later call will be authorized. The exercise should expose missing permissions while time remains to correct them deliberately.

Bind validation to the correct resource

Restore testing can start a workflow when a job reaches COMPLETED. Filter events to intended scope, such as resource type and plan ARN, and use the created resource’s identity. A fixed production endpoint can respond successfully while the restore is unusable. Define read, consistency, and functional checks before publishing the result. An already set validation status cannot be changed; do not send SUCCESSFUL as a placeholder. Retain logs linking job, resource, and checks. Validation completion or window expiry initiates cleanup, which depends on resource type and may not be immediate. Collect evidence and track outstanding cleanup.

Respect retention and restore semantics

Vault Lock governance permits removal by authorized users; compliance makes the lock immutable after the change period. That decision requires understanding retention and configuration correction. New jobs with incompatible retention can fail rather than receiving automatic adjustment. Check outcomes rather than merely plan existence. For an RDS instance, point-in-time restore creates another instance without modifying the source. Review configuration and access and plan client cutover. Available state can coexist with background block loading and still-settling performance. Validate required load before declaring recovery of capacity supporting close.

Calculate the critical path and accept the outcome

The local model adds eight minutes of detection and decision, the longer of two parallel branches lasting eighteen and twelve minutes, and seven minutes of validation: thirty-three minutes. Do not add parallel work as sequential or omit steps within the agreed interval. The code declares dependencies and rejects cycles; it does not perform a real restore. Use it to discuss where improvement reduces the total: accelerating the twelve-minute branch changes nothing while the eighteen-minute branch dominates. During rehearsal, replace estimates with observations and document limitations. Final acceptance combines time, data loss, functionality, capacity, and outstanding work with owners.

Control impact during fault rehearsal

A FIS rehearsal should start with defined hypothesis, targets, steady state, and impact boundaries. Connect a stop condition to a CloudWatch alarm representing service-relevant degradation. A Running experiment does not prove customer experience remains acceptable. Stopping ends the experiment and does not allow resuming that execution; outstanding post actions are handled before stopping, but the team remains responsible for checking resulting functional state and required recovery work. Plan who decides, who observes, and who may perform recovery. Do not change the limit during degradation merely to keep the rehearsal running. Record threshold-breach time and observed response to evaluate the design against its stated hypothesis.

Accept only coverage actually rehearsed

Target preview with skip-all helps inspect selection, logging, and certain account configurations but does not inject faults or validate every action permission. Targets may change between preview and execution through resource changes or sampling. With emptyTargetResolutionMode=skip, actions without filter-resolved resources may be skipped; that does not demonstrate intended-resource recovery. In reporting, distinguish what was planned, selected, actually affected, and validated. The local model below rejects a complete-rehearsal claim when only preview exists, targets remain unexercised, or functional validation is absent. It neither calls FIS nor injects faults. Use it to prepare project-evidence review before RUN handover, keeping incomplete coverage visible rather than counting template steps as executed tests.

# Original local dependency model, no AWS calls or resource changes.
def completion_times(tasks):
 done, visiting = {}, set
 def finish(name):
 if name in done:
 return done[name]
 if name in visiting:
 raise ValueError("Dependency cycle")
 if name not in tasks:
 raise ValueError("Unknown dependency")
 duration, dependencies = tasks[name]
 if type(duration) is not int or duration < 0:
 raise ValueError("Invalid duration")
 visiting.add(name)
 done[name] = max((finish(d) for d in dependencies), default=0) + duration
 visiting.remove(name)
 return done[name]
 for name in tasks:
 finish(name)
 return done

plan = {"decision": (8, []), "database": (18, ["decision"]),
 "compute": (12, ["decision"]), "validation": (7, ["database", "compute"])}
assert completion_times(plan)["validation"] == 33
assert completion_times({**plan, "compute": (5, ["decision"])})["validation"] == 33
assert completion_times({**plan, "database": (10, ["decision"])})["validation"] == 27
try:
 completion_times({"a": (1, ["b"]), "b": (1, ["a"])})
 raise AssertionError("Cycle accepted")
except ValueError:
 pass

# Original local coverage exercise; not an FIS executor or recovery guarantee.
def rehearsal_evidence(expected, exercised, validated, preview_only):
 if not expected:
 return "scope missing"
 if preview_only:
 return "configuration evidence only"
 if expected - exercised:
 return "fault coverage incomplete"
 if expected - validated:
 return "functional validation incomplete"
 return "recorded scope exercised and validated; review residual risk"

scope = {"worker-a", "worker-b"}
assert rehearsal_evidence(scope, set, set, True) == "configuration evidence only"
assert rehearsal_evidence(scope, {"worker-a"}, scope, False) == "fault coverage incomplete"
assert rehearsal_evidence(scope, scope, {"worker-a"}, False) == "functional validation incomplete"
assert rehearsal_evidence(scope, scope, scope, False).startswith("recorded scope")
assert rehearsal_evidence(set, set, set, False) == "scope missing"
IN PRACTICE

The RDS job ends in 18 minutes, but service returns only after connection cutover, data validation, and capacity recovery. Reporting retains both times and compares the second with the agreed objective.

Common pitfalls

Validating production instead of the restore; publishing success before checks; assuming PITR changes the source; timing only the job; ignoring reactivated integrations.

Related topics: Readiness, scaling, and controlled exit

Take this idea with you

Demonstrated recovery connects the data point, resource, dependencies, timing, and functional acceptance. A job state alone does not cover that set.

Create account

Reference: AWS Backup restore testing validation · DOP-C02

AWS is a trademark of Amazon.com, Inc. or its affiliates. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by AWS. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.