Concept and mechanism
Systems Manager depends on a managed node reachable through its agent with suitable credentials and connectivity. On private EC2, investigate SSM Agent state, role, resolution, and outbound access to required endpoints. Opening SSH to the Internet does not fix the agent management path. An operational runbook must identify targets precisely: an overly broad tag can turn maintenance of one service into action across several. Record parameters and runbook version so outcomes can be tied to actual execution. Privileges should match delegated actions rather than every task the team might perform.
Guided application
Before a disruptive change, an aws:approve step can enforce required approval with identified approvers and correct permissions. For multiple targets, configure concurrency and error thresholds consistent with acceptable exposure. Reaching a threshold stops new dispatch but does not instantly undo in-flight executions. Monitor every outcome and retain a defined recovery procedure. A small initial batch can reveal a wrong assumption before the whole perimeter is affected. Reporting to the PM should distinguish started, completed, failed, and unexecuted targets; a stopped overall execution does not mean every target remained unchanged.
For batch patching, confirm two initial targets and service behavior after restart before increasing concurrency through the approved plan.
Common pitfalls
Unreviewed tags; approval only in a comment; error thresholds treated as undo; forgetting in-flight actions after stopping.
Related topics: Security in operations and recovery · Networking, DNS, and content delivery
Safe automation requires controlling scope and observing each outcome.
Reference: Automation rate controls · SOA-C03; exam guide 1.1