Choose the mechanism from intent
In a fictional middleware fleet, three tasks look similar: maintaining agent configuration, applying a one-off correction, and installing patches in a window. First define desired state, scope, and frequency. State Manager associates configuration with targets and scheduling; Run Command dispatches commands and returns invocation outcomes; Automation links steps and decisions. Maintenance Windows organizes execution within defined periods. Combining services does not remove the need for an outcome contract. For each task, state document version, parameters, identity, expected targets, preconditions, and subsequent validation. A command finishing with exit code zero may still fail to demonstrate recovery of the business service. Define the functional observation required before declaring success.
Treat an association as a change
Creating an association is not necessarily inert preparation for Sunday. Immediate execution is the default; with a cron schedule, ApplyOnlyAtCronInterval can restrict that behavior. The parameter does not support rate expressions. Document version also matters: DEFAULT and LATEST are moving references, and same-account changes can trigger reapplication. Decide whether the change needs a pinned numeric version and explicit promotion. Before saving the association, confirm actual targets and the authorized window. If some nodes already changed, correcting the schedule does not automatically recover them. Contain further changes, collect partial inventory, and decide recovery according to observed impact. Keep document identity and execution identity distinct in the report.
Observe dispatch, execution, and effect
Run Command reports plugin, invocation, and aggregate command states. Success with zero selected nodes does not fulfill a request to configure twelve servers. Compare expected targets with outcomes and investigate Delivery Timed Out at the delivery layer before blaming script logic. The API is eventually consistent; a read immediately after dispatch may not yet show the invocation. Poll the same identifier with bounded waiting instead of repeatedly launching work that may already be running. The error threshold prevents further dispatch once exceeded, but dispatched invocations can continue. Record final effects and functional health rather than request acceptance alone. Preserve enough detail to retry only the work that still needs attention.
Build automation that preserves failure
Automation may need to continue after failure to collect logs or release temporary resources. If the change is critical, express isCritical and onFailure flow explicitly; do not let successful cleanup conceal the main failure. In executeScript, calls subject to throttling need retry handling in the script itself, with limits and retry-safe operations. The action has a maximum duration, so unbounded loops are not a recovery strategy. Distinguish transient failures from authorization or input errors requiring correction. In the operational exercise, ask a colleague to explain which effects may already exist before repeating a step and where proof is retained. Recovery must preserve both useful diagnostics and an honest final outcome.
Separate installation, activation, and continuity
Start patching with Scan when the objective is to assess compliance without installing. Before Install, confirm baseline, dependencies, alternative capacity, and return-to-service criteria. Outside Maintenance Windows, a shared Snapshot-ID for the same baseline and operation helps maintain a consistent approved set across nodes. NoReboot controls operating-system reboot, but packages can restart services. Do not treat the option as an availability guarantee. InstalledPendingReboot leaves an action outstanding; schedule reboot and another Scan rather than editing tracking files to produce a green dashboard. Acceptance includes active updates, service health, and evidence for every intended node. A technically completed installation can still leave operational work unfinished.
Exercise: review window headroom
The local model calculates the new-start deadline from UTC start, duration, and cutoff. It also compares a task estimate with remaining window time as an additional internal review. It does not reproduce the AWS scheduler or promise to cancel processes at window end. Try a task starting before cutoff but taking too long, a proposal after cutoff, and an estimate that fits. In the real plan, make ScheduleTimezone, upcoming executions, cohorts, recovery margin, and stop-decision ownership explicit. An estimate is not a guarantee: overrun requires communication and control of in-progress state rather than declaring success merely because the window ended. Preserve separate evidence for scheduling and operational completion.
# Original local planning review, not an AWS scheduler or cancellation guarantee.
from datetime import datetime, timedelta, timezone
def review_start(start, duration_hours, cutoff_hours, proposed, estimate_minutes):
if start.tzinfo is None or proposed.tzinfo is None:
raise ValueError("timezone-aware timestamps required")
if not 0 <= cutoff_hours < duration_hours or estimate_minutes <= 0:
raise ValueError("invalid planning inputs")
end = start + timedelta(hours=duration_hours)
deadline = end - timedelta(hours=cutoff_hours)
# This educational policy treats the exact cutoff instant as closed.
if not start <= proposed < deadline:
return "outside local start interval"
if proposed + timedelta(minutes=estimate_minutes) > end:
return "estimate exceeds window; review scope or timing"
return "fits estimate; authorization and recovery margin still required"
start = datetime(2026, 10, 11, 2, tzinfo=timezone.utc)
assert review_start(start, 4, 1, start + timedelta(hours=2), 60).startswith("fits")
assert review_start(start, 4, 1, start + timedelta(hours=2, minutes=50), 90).startswith("estimate exceeds")
assert review_start(start, 4, 1, start + timedelta(hours=3, minutes=1), 10).startswith("outside")
assert review_start(start, 4, 1, start - timedelta(minutes=1), 10).startswith("outside")
try:
review_start(start.replace(tzinfo=None), 4, 1, start, 10)
except ValueError:
pass
else:
raise AssertionError("naive timestamp accepted")
Creating the association on Friday started changes intended for Sunday; containing new starts does not by itself recover affected nodes.
Common pitfalls
Cron as an initial-wait guarantee; Success without targets; resending after early reads; Continue hiding failure; NoReboot as zero impact.
Related topics: StackSets: scope and destination outcomes
Maintenance ends when intended state and functional health are demonstrated, not merely when a call or window ends.
Reference: State Manager execution behavior · DOP-C02