Draw scope before rollout
A fictional team manages connectivity and logging across development and production accounts. The request says all accounts, but the design contains only an OU name. Before execution, turn that phrase into an inventory of accounts, Regions, expected resources, and approved exceptions. StackSets groups stacks based on a common template with parameters that can vary. Identify the Region where the StackSet is administered and target Regions; these are separate decisions. The service-managed model integrates Organizations and creates required roles, while self-managed requires explicit trust relationships. This choice affects onboarding, authority, and coverage. Acceptance of an operation in the central account does not prove deployment at its destinations.
Prepare account entry and exit
With service-managed permissions, the management account does not receive stacks through this mechanism. Coverage planning must address that requirement separately. With automatic deployments, adding an account can create resources using StackSet defaults without inheriting older accounts’ overrides. Filters used in an initial operation also do not guarantee exclusion of future accounts from automatic deployment. In this lesson’s case, internal policy requires 90-day retention while the default remains 30. Onboarding must correct that mismatch before entry. When an account leaves the OU, retaining stacks preserves resources outside the StackSet; assign ownership, cost, maintenance, and eventual retirement instead of declaring those resources gone.
Separate standard, exception, and drift
Record the reason and expiry of parameter exceptions. Updating a default does not automatically remove an existing override. Drift answers a different question: do resources differ from the stack’s expected state? Updating an individual stack through CloudFormation can change its template without producing resource drift, even though it no longer represents the intended central standard. Therefore compare effective templates and parameters when that consistency belongs to the requirement. Successful DETECT_DRIFT can discover discrepancies; observation success is not repair. Reconciliation must decide whether to correct the resource or update approved intent while preserving evidence and dependencies. Do not silently redefine the standard to match an unexplained difference.
Plan exposure by account and Region
Concurrency and tolerance describe execution behavior rather than the business impact accepted without analysis. Strict mode reduces concurrency as failures occur; soft mode maintains its intended window and can finish with additional failures from dispatched work. The arithmetic exercise uses the documented planning rule rather than a ready-to-submit API command. Also decide whether Regions should progress sequentially or in parallel. RegionOrder does not create a human gate: if the second Region requires approval after observing the first, divide the change into controlled operations. Manual-operation preferences do not automatically apply to AutoDeployment onboarding. Review those two paths separately and communicate the actual exposure before execution.
Contain and recover partial state
Acceptance of a stop request does not mean every account stopped or returned to its prior configuration. StopStackSetOperation cancels unstarted deployments and waits for work already running. Keep the operation ID and follow each outcome until completed, failed, and unstarted destinations can be separated. With ManagedExecution, requests can queue, including apparently non-conflicting requests when work is already active. Do not duplicate them because results are not immediate. The sponsor needs a factual account: which destinations received the change, which are still changing, which are missing, and who decides recovery. Sending the reverse change indiscriminately can overlap operations acting on different states.
Exercise: close coverage per destination
The local model compares an expected set of account/Region pairs with synthetic outcomes and the intended revision. It identifies missing destinations, failures, wrong revisions, and results outside intended scope. It does not call AWS, interpret every CloudFormation state, or resolve authorization. Execute the examples, add a new account, and remove a Region from reporting. Explain how an aggregate SUCCEEDED operation can leave the coverage objective unmet when failure tolerance exists. At change closure, present the matrix, exceptions, and next action per row. Retain evidence so another person can reconstruct the decision without having watched execution. Coverage is an explicit comparison, not a conclusion drawn from one status label.
# Original local inventory review, not a CloudFormation state machine.
def review_rollout(expected, observed, revision):
gaps = []
for target in sorted(expected):
row = observed.get(target)
if row is None:
gaps.append((target, "missing result"))
elif row["status"]!= "SUCCEEDED":
gaps.append((target, "not completed successfully"))
elif row["revision"]!= revision:
gaps.append((target, "wrong revision"))
return {"gaps": gaps, "unexpected": sorted(set(observed) - expected)}
expected = {("account-a", "eu-west-1"), ("account-b", "eu-west-1")}
good = {k: {"status": "SUCCEEDED", "revision": "network-7"} for k in expected}
assert review_rollout(expected, good, "network-7") == {"gaps": [], "unexpected": []}
assert len(review_rollout(expected, {}, "network-7")["gaps"]) == 2
bad = {**good, ("account-b", "eu-west-1"): {"status": "RUNNING", "revision": "network-7"}}
assert review_rollout(expected, bad, "network-7")["gaps"][0][1] == "not completed successfully"
assert len(review_rollout(expected, good, "network-8")["gaps"]) == 2
extra = {**good, ("account-c", "eu-west-1"): {"status": "SUCCEEDED", "revision": "network-7"}}
assert review_rollout(expected, extra, "network-7")["unexpected"] == [("account-c", "eu-west-1")]
A new account receives 30-day retention when policy requires 90; technical onboarding success does not prove parameter compliance.
Common pitfalls
Assume inherited overrides; omit the management account; treat drift as universal comparison; assume rollback on stop; use SUCCEEDED as coverage proof.
Related topics: Configuration, patching, and operational windows
Actual scope and per-destination outcomes must match approved intent, including new accounts and persistent exceptions.
Reference: StackSets concepts · DOP-C02