1. Define the service that must return
Begin with the outcome the business needs to recover. In a fictional funds application, reading positions does not demonstrate the ability to receive a file, validate movements, and reconcile processing. Define operations, dependencies, and acceptance criteria before rehearsing. If the job takes 22 minutes and required sequential checks another 18, observed time is 40 minutes for that scope. A 30-minute objective was not achieved. This calculation is an operational convention for the exercise, not a universal AWS metric. Also record the recovery-point timestamp and its difference from expected data; finishing quickly does not repair excessive data loss. Name who can accept the service or activate approved contingency.
2. Select resources and recoverable points
A restore testing plan needs resource selection; scheduling does not perform that assignment. With tag-based selection, tags on the latest recovery point help select the resource. Subsequently choosing a random point does not require that point to carry the same tags. A resource without an eligible point in the window did not participate in a successful restore: it was left out. Consider actual execution when interpreting the selection window. A delayed execution can put a point outside the interval even if it appeared eligible at scheduled time. For the manager, provide intended, selected, excluded, and restored resource lists. Require an explanation for each difference before presenting the rehearsal as complete application coverage.
3. Diagnose context and permissions
The rehearsal environment may lack the default VPC expected by inferred metadata. Inspect expected metadata, compare it with values used by the job, and specify an appropriate rehearsal subnet. Do not repeatedly submit the same request while expecting a different network. If AWS Backup cannot assume the role, start with its trust relationship; adding actions to a permissions policy does not repair who can assume the identity. For encrypted backups, check permissions and key policies relevant to both backup and restored resource. Operator console access does not establish role access. In the runbook, connect each error message to the responsible team and the evidence needed to confirm repair.
4. Validate before reporting success
A restore reaching COMPLETED can trigger a custom workflow through EventBridge. Have that workflow run agreed checks and only then publish its result using PutRestoreValidationResult. A recorded validation status cannot be changed; SUCCESSFUL is not a provisional marker while reconciliation runs. Retain logs and identifiers that explain functional failure. Restore testing resources are temporary and enter cleanup after validation or window expiry. For resources receiving the awsbackup-restore-test tag, removing it can prevent automatic cleanup. The local exercise below evaluates fictional evidence and never calls the API. If a business check fails, keep acceptance pending and investigate without treating resource existence as proof of readiness.
5. Review retention before irreversibility
Governance permits Vault Lock removal with sufficient authorization. Under compliance, use LockDate to identify when changing the lock stops being possible; Locked=true alone does not describe that phase. Review retention and costs with FinOps before that point, especially recovery points set to Always. Retention bounds reject incompatible new jobs and do not rewrite the lifecycle of older points. In the change exercise, a seven-day copy into a vault with a thirty-day minimum fails; it is not automatically assigned thirty days. Protection from deletion also does not prove immediate restore capability: a disabled AMI can prevent the operation. Acceptance planning should check retention, accessibility, and functionality as separate evidence.
6. Rehearse recovery across accounts
Sharing a logically air-gapped vault uses individual recipient accounts, not an OU. The recipient can inspect and restore shared points but cannot create copies of them; additional IAM permissions do not remove that limitation. Current encryption choices allow the default AWS-owned key or a customer managed key at creation, without subsequently changing the vault key. Turn these constraints into design questions: who retains the copy, who restores, and who demonstrates dependency access? In an international rehearsal, distribute responsibilities and shift-transition criteria. Summarize results by service, recovered point, validated operations, observed time, and gaps. These are original APS examples, not the internal procedures of any bank.
# Original fictional acceptance model, not the AWS validation API.
# No credentials, network calls, or changes to any backup.
def acceptance(job_complete, checks, restore_minutes, validation_minutes, objective):
if not job_complete:
return "restore-incomplete"
if not checks or not all(checks.values):
return "functional-evidence-incomplete"
if restore_minutes + validation_minutes > objective:
return "time-objective-missed"
return "accept-within-stated-scope"
assert acceptance(False, {}, 0, 0, 30) == "restore-incomplete"
assert acceptance(True, {}, 22, 0, 30) == "functional-evidence-incomplete"
assert acceptance(True, {"read": True, "reconcile": False}, 22, 18, 30) == "functional-evidence-incomplete"
assert acceptance(True, {"reconcile": True}, 22, 18, 30) == "time-objective-missed"
assert acceptance(True, {"reconcile": True}, 22, 18, 40) == "accept-within-stated-scope"
print("five recovery-evidence cases passed; no validation result was published")
A restore completes but funds reconciliation fails. The team keeps acceptance pending and communicates impact and contingency.
Common pitfalls
COMPLETED as available service; excluded points as tested; Locked as proof of grace-time expiry; extra permissions as a universal fix.
Related topics: Incident response · Acceptance and RUN handover
A restore supports acceptance only within the demonstrated functional and timing scope.
Reference: Restore testing · SCS-C03