Identify the operation before the policy
An infrastructure change involving data requires identifying whether it updates the same resource, replaces it, or removes it from management. A template’s logical identifier can stay unchanged while the physical instance changes. In the fictional case, a reporting database is replaced during a release. DeletionPolicy: Retain alone does not govern the old instance along that path; UpdateReplacePolicy is relevant. Also review references, data migration, and service continuity. Successful resource creation does not establish that required state is present. Change review should describe sequence, preservation evidence, and the criteria for accepting or recovering each step.
Assign ownership to preserved resources
Retention avoids deletion on that path but does not turn the old copy into an automatically managed or free database. A resource retained after replacement leaves stack scope; it still needs access control, protection, cost ownership, and future deletion. Snapshot is an UpdateReplacePolicy option only where the type supports it; the reference warns of fallback to Delete otherwise. RetainExceptOnCreate distinguishes rollback of the resource-creating operation from later deletion. Choose based on requirements rather than a reassuring name. Handover should include physical identifiers, owners, review dates, and evidence needed before retained resources or snapshots can be removed.
Combine protections with different scopes
Termination protection protects attempts to delete the stack; it is not a universal prohibition on destructive updates. Updating the parent can remove a nested stack. Stack policy controls CloudFormation update actions on resources without replacing IAM controls on direct service calls. With overlapping Allow and Deny, explicit denial takes precedence. Current guidance also warns that exclusion through NotResource with Allow does not reliably protect because logical ID and type are evaluated; use explicit protection and test it. Review should establish which operation is prevented, on which path, and who can change that protection.
Recover control without hiding discrepancy
UPDATE_ROLLBACK_FAILED is not a state where indiscriminately repeating a normal update solves the issue. Read events, identify the cause, and prepare ContinueUpdateRollback when conditions allow. Skipping eligible resources through ResourcesToSkip can unblock recovery but leaves an obligation to reconcile differences from the template. Use the minimum necessary set and record the exception. Application availability does not establish that the next stack update will be safe. After importing existing resources, check drift within supported scope and reconcile intent with reality. In both cases, distinguish recovered administrative state from configuration that has actually been understood and validated.
Handle dependencies and protocols
Exports used through Fn::ImportValue create dependencies restricting value changes or deletion of the exporting stack. Plan consumer migration before retiring that source. A Secrets Manager dynamic reference also has a lifecycle: changing only the secret does not make CloudFormation fetch it automatically; validate rotation strategy and the consuming service. For custom resources, PhysicalResourceId returned on Update distinguishes update from replacement. Generating a random identifier can trigger Delete against the previous object. The handler must also respond to ResponseURL; finishing the function without error does not itself report success to the orchestrator. Reconcile effects before retrying requests whose outcome became uncertain.
Handover and decommission exercise
The local model reviews a retained old resource without deleting anything. It returns specific reasons when ownership, reconciliation, approval, absence of consumers, or retention clearance is missing. Four cases check individual blockers, simultaneous blockers, and eligibility for a planned change. The model trusts supplied values: True is not proof of real authorization or verification of every dependency. In the team exercise, associate each value with evidence, an owner, and a verification date. The PM brings APS, application teams, and FinOps together to accept necessary coexistence and correct the forecast. The final summary should let another shift identify what can change, what must be preserved, and which evidence remains missing.
def retirement_review(resource):
checks = {
"missing-owner": bool(resource.get("owner")),
"reconciliation-pending": resource.get("reconciled") is True,
"approval-pending": resource.get("approved") is True,
"active-consumers": resource.get("consumers") == 0,
"retention-pending": resource.get("retention_released") is True,
}
reasons = [name for name, passed in checks.items if not passed]
return {"eligible_for_planned_change": not reasons, "blocking_reasons": reasons}
ready = dict(owner="APS", reconciled=True, approved=True,
consumers=0, retention_released=True)
cases = [
({**ready, "reconciled": False}, ["reconciliation-pending"]),
({**ready, "owner": "", "consumers": 2}, ["missing-owner", "active-consumers"]),
({**ready, "approved": False, "retention_released": False},
["approval-pending", "retention-pending"]),
(ready, []),
]
for resource, expected in cases:
result = retirement_review(resource)
assert result["blocking_reasons"] == expected
assert result["eligible_for_planned_change"] == (not expected)
# Original local review model. No deletion, AWS calls, live inventory or policy discovery.
# It trusts supplied evidence: a True value is not proof of real authorization or retention clearance.
Fictional case: the stack recovers after a skip, but one resource differs from the template; RUN handover retains the exception with ownership and planned correction.
Common pitfalls
Treating Retain as backup or zero cost; relying on termination protection for every update; forgetting skipped resources; confusing function success with protocol response.
Related topics: Pipeline concurrency, gates, and recovery
A change is operationally closed only when state, evidence, ownership, and disposition of old resources are reconciled.
Reference: CloudFormation DeletionPolicy · DOP-C02