1. Write scope as a requirement
A central policy does not automatically protect every resource the team has in mind. In Firewall Manager, account selection combines with resource types and tags. ALL and ANY represent different choices; tag matching uses key and value, and an omitted value is an empty string. Translate the business request into examples inside and outside scope. Include a new account, a missing tag, and a resource changing OU or tag value. Report expected and observed populations. This comparison helps identify accidental exclusions before the team uses a compliance percentage as proof of complete coverage.
2. Define who maintains state
If Terraform creates a rule and Firewall Manager removes it, two tools enforce different intentions. Increasing retries does not resolve ownership. Align configuration or define mutually exclusive scopes, using bounded approved exceptions where necessary. Leaving scope also does not mean every protection disappears: by default, an ALB can retain its web ACL when automatic removal is not selected. The migration plan must state which association remains, who manages it, and how state is confirmed after both tools run their cycles. During a short window, a defined rollback or bounded correction is preferable to a global shutdown without inventory. Record operational debt when handover is phased.
3. Read the SCP path
For an action to pass the SCP boundary, it needs Allow at every level along root, OUs, and account, without an applicable Deny. Multiple Allows can combine within one level; a broad Allow can defeat the intent of a shorter allowlist at that level. Across levels, a lower Allow neither fills an upper omission nor bypasses inherited Deny. Draw the account effective path before changing IAM. A delegated-administrator member account remains subject to applicable SCPs; delegation does not make it the management account. Rehearsal should include operational dependencies such as evidence collection and recovery so a security policy does not interrupt response capability itself.
4. Maintain assessment and evidence
A conformance pack is service-managed. Editing its underlying stack directly can create drift and rules that are difficult to remove; use the pack mechanism and investigate failures through status and events. Deleting the pack removes rules, remediations, and results, so required evidence must be preserved beforehand. Deletion is asynchronous. If you suspend ResourceCompliance recording to reduce CIs during that operation, document the history gap and track restoration after completion. In reporting, the score is a proportion of rule-resource combinations: 18 out of 20 and 90 out of 100 are both 90%, but represent different populations. Report denominator, scope, and concrete failures without presenting the number as legal certification.
5. Control new dispatch without inventing rollback
In incident response, Systems Manager Automation can apply a runbook to multiple targets. MaxConcurrency limits simultaneous work and MaxErrors controls when new targets stop launching after errors. With zero, the first received failure stops new dispatch; it does not prevent starting the first target. Executions already running can continue and fail. This lesson local model separates fleet completion from stopped dispatch to make that distinction visible. It does not simulate the AWS scheduler or calculate error percentages. In production, inspect actual states before repeating, cancelling, or declaring recovery. An observation taking longer than expected is not evidence that the operation ended.
6. Pin execution and prepare closure
A numeric DocumentVersion allows execution of the rehearsed revision. DEFAULT and LATEST can point to a different revision when the window begins. Also record parameters, targets, role, and per-location controls: TargetsMaxConcurrency and TargetsMaxErrors take precedence over general values when supplied. Closure needs per-target evidence, produced effects, failures, recovery, and operations still live. Communicate the difference between stopping new dispatch and completing containment to management. When targets remain pending, retain an owner and the next status checkpoint. End the retrospective with a concrete runbook or acceptance-criteria change and a rehearsal date. The goal is for the next team to operate the control without depending on informal memory.
# Original fictional closure model, not an AWS scheduler or cancellation tool.
# No credentials, network calls, or AWS changes.
TERMINAL = {'Success', 'Failed', 'Cancelled', 'TimedOut'}
def closed(states):
return bool(states) and all(s in TERMINAL for s in states)
assert closed(['Success','Failed']) is True
assert closed(['Success','Running']) is False
assert closed(['Pending']) is False
assert closed([]) is False
assert closed(['Cancelled','TimedOut']) is True
print('five closure-state cases passed; no automation was cancelled')
The dashboard window ends, but three targets are still executing containment actions.
Common pitfalls
Lower Allow as override; empty tags as wildcard; deleting a pack as archiving; stopped dispatch as rollback.
Related topics: Evidence and operational acceptance · Security and continuity
Accept a policy only when scope, effects, and operational limits are demonstrated.
Reference: Firewall Manager policy scope · SCS-C03