← AWS CloudOps Engineer: operations and recovery
04 / 7 · 30 MIN

Stacks, images, and deployments

Keep configuration reproducible and recovery viable.

Concept and mechanism

An urgent change outside CloudFormation can resolve an incident while leaving drift. Detection compares supported resources and properties; NOT_CHECKED does not prove conformity and aggregate IN_SYNC does not cover everything. Decide whether the live change should remain or be reversed using requirements and incident authorization. Then reconcile the template and rehearse the update. Blindly copying live state can perpetuate an inappropriate exception; blindly restoring the template can reintroduce failure. When a stack fails, preserve events, identify the first relevant cause, and confirm the actual action and identity before correcting permissions or dependencies.

Guided application

An image lifecycle includes future references. An AMI may not have been used recently yet still be needed by a launch template to replace instances. Before retirement, update references and demonstrate a successful fresh launch. Existing instances do not prove that capability. For an ECS rolling-deployment service, the circuit breaker can detect failures and support rollback to a completed deployment. If none ever reached COMPLETED, that reference is missing. APS should know failure signals, approved image, and corrective steps without assuming enabling rollback automatically creates a healthy version.

IN PRACTICE

A scale-out rehearsal before AMI retirement reveals stale references missed by observing only running instances.

Common pitfalls

Drift treated as automatic remediation; deleting stacks before reading events; old image treated as unreferenced; rollback without a valid target.

Related topics: Controlled operational automation · Security in operations and recovery

Take this idea with you

Configuration must explain how to create and recover future resources.

Create account

Reference: CloudFormation drift detection · SOA-C03; exam guide 1.1