1. Know who the mechanism covers
Start with a list of accounts, OUs, Regions, and deployment channels. An account moved into a registered OU might not yet be enrolled in Control Tower. OU preventive controls cover unenrolled accounts, but detective and proactive controls do not have the same coverage. A proactive control uses CloudFormation hooks; a template result does not demonstrate blocking of a direct call that does not traverse that mechanism. In a fictional acquisition, governance should record pending enrollment and compensating controls before declaring the account integrated. For each requirement, write down the mechanism and its actual population. This matrix exposes gaps that the account position in the hierarchy does not reveal.
2. Treat drift as a scoped comparison
CloudFormation drift detection compares supported configurations with the managed definition. NOT_CHECKED should remain an evaluation gap, and an IN_SYNC stack can contain unsupported resources. Nested stacks need their own operation. Explicitly define in the template or parameters a supported property you intend to track; relying only on the service default can leave it outside comparison. In the committee report, state the compared population, excluded resources, and observation time. The local exercise separates these states instead of converting every non-red result to success. The model does not query CloudFormation or know actual configuration. It supports discussion of missing evidence before accepting a resource set.
3. Compare the central reference too
A StackSet stack can be updated directly through CloudFormation using another template. Resources can remain equal to their local definition and show no drift even though the central baseline differs. Add a comparison of approved versions and parameters to demonstrate uniformity. FinOps encounters a similar population issue: a favorable tag report does not necessarily cover resources that have never had a user tag. Reconcile report population with suitable inventory. For required-key reporting, documentation identifies an account-level Resource Groups console limitation and points to the organization-wide report. Select the evidence mode that answers the requirement; an available dashboard is not automatically a sufficient dashboard.
4. Choose policies and rehearse impact
An RCP limits permissions on covered resources but does not restrict service-linked role calls. Before expanding a policy, include legitimate integrations and requests that must be denied in a scoped pilot. An isolated internal test does not demonstrate compatibility for external partners. For a supported service attribute, a declarative policy can enforce configuration in the control plane instead of maintaining an API action list. Effective policy and scope still need review. The retirement plan also matters: documentation describes attribute reversion to its previous state when a declarative policy is detached. Identify that state and the replacement control before treating detachment as a simple administrative task.
5. Remediate the state that exists now
Config automatic remediation can start from an old snapshot. An operator may have fixed the resource between evaluation and execution. The operational recommendation is to query current state, make the action safe when the fix already exists, and record the decision not to repeat an unnecessary change. This recommendation is a design inference, not a guarantee that every runbook already implements it. During a race at daily close, additional retries can multiply restarts without adding value. To diagnose failures, query execution status with steps, times, and errors. Distinguish intended configuration, historical evaluation, and observed outcome; these records answer different questions from a manager or auditor.
6. Present a decision with clear limits
Build the acceptance package with requirement, approved reference, covered accounts and resources, observation time, result, and gaps. Assign an owner and next action to each relevant gap. A subset can be accepted when the decision defines scope and retains follow-up for remaining elements; that does not authorize declaring the entire platform compliant. At RUN handover, confirm who observes controls, who can change policies, and who responds to failures outside working hours. At resource retirement, include policy dependencies, evidence to retain, and final-state conditions. This lesson summary is to separate intent, application, observation, and decision. That discipline communicates real progress without turning no alert into a broader conclusion than the data supports.
# Original evidence inventory exercise, not CloudFormation drift detection.
# No credentials, network calls, or AWS changes.
def classify(row):
if row["status"] == "NOT_CHECKED":
return "needs-evidence"
if row["status"] == "DRIFTED":
return "deviation"
if row["status"]!= "IN_SYNC":
return "unknown"
if row["template"]!= row["approved_template"]:
return "baseline-mismatch"
return "matches-recorded-scope"
base = dict(status="IN_SYNC", template="v4", approved_template="v4")
assert classify(base) == "matches-recorded-scope"
assert classify({**base, "status": "NOT_CHECKED"}) == "needs-evidence"
assert classify({**base, "status": "DRIFTED"}) == "deviation"
assert classify({**base, "status": "FAILED"}) == "unknown"
assert classify({**base, "template": "local-v5"}) == "baseline-mismatch"
print("five evidence-scope cases passed; scope is not whole-platform compliance")
A stack without drift uses a divergent local reference; acceptance requires comparing the central baseline and addressing unevaluated resources.
Common pitfalls
OU membership as enrollment; no drift as uniformity; tag report as complete inventory; stale snapshot as current state.
Related topics: Telemetry and investigation scope · Operational handover and acceptance criteria
A conclusion must state the reference and population against which it was reached.
Reference: Control Tower terminology · SCS-C03