← AWS DevOps Engineer Professional: operations and delivery
22 / 24 · 80 MIN

Recovery workflows and preconditions

Design recovery that preserves context, bounds repetition, and confirms actual state before and after resource changes.

Define the recovery contract

A fictional team automates connectivity recovery after an alert. The flow identifies the service, waits for approval, changes a route, and validates a functional operation. Before choosing services, write down the intended outcome, deadline, authority, permitted effects, and closure evidence. Human waiting may exceed an Express workflow’s limit; Standard supports a durable design with appropriate integration patterns. That does not remove responsibility for controlling external changes. Execution history should reveal which version made the decision, against which resource, and with what input. Also define what happens when approval never arrives or validation cannot observe the result. These unresolved paths belong in the design before release.

Separate request, execution, and effect

A client can lose the start response while execution is already running. Preserve identity and observe the same work before launching another execution. Standard StartExecution idempotency has specific name, input, and state conditions; it should not be generalized to Express or the whole business operation. Inside a workflow, Request Response waits for the API response and can advance before a job finishes. Use a pattern matching the need to wait, such as.sync for supported services. If execution is stopped, inspect the destination too: cancelling external work is an attempt that can fail. A terminal orchestrator state does not automatically close investigation of effects.

Bound retries and retain context

Classify failure before choosing repetition. Transient unavailability may justify backoff, jitter, and limits; an invalid input contract needs correction. Count the initial call in addition to authorized retries and consider total time rather than individual delay alone. Error handling has boundaries: States.ALL does not resolve States.Runtime caused by invalid data processing. On the diagnostic path, retain resource identity and required context. In a JSONPath workflow, ResultPath in Catch can combine the error with input instead of replacing it. Record the main failure even when diagnostic collection succeeds. Cleanup must not turn a failed change into apparent success, and every retry still needs a justified effect model.

Choose how to resume partial work

A Task may apply a change and fail before communicating success. Redrive returns to the unsuccessful step while preserving earlier results and the original definition. Publishing a fixed definition or moving an alias does not automatically make that redrive use new logic. If the definition fix is necessary, assess a new execution and first reconcile what already happened. In the routing example, inspect configuration and test traffic before repeating the write. Use stable operation identity where the destination supports it and define compensation where repetition is unsafe. Technical redrive eligibility is not change authorization or proof that partial effects are absent from the target system.

Validate remediation preconditions

AWS Config can start correction from a snapshot that no longer represents the resource. A legitimate manual intervention may have fixed the property meanwhile. The runbook should read current state, confirm identity and scope, and avoid unnecessary change when the goal is already satisfied. Explicitly configure the automation role and permissions required by the action. Concurrency and error limits help contain exposure but replace neither a correct precondition nor subsequent validation. Test compliant, noncompliant, unavailable, and concurrently changing states. If the read fails, do not assume noncompliance to justify a write. Record the decision not to change as an observable, explainable result with evidence supporting it.

Exercise: decide before writing

The local model separates three decisions: authorization is absent, current evidence is missing, or a difference justifies preparing a change. It does not call AWS, guarantee mutual exclusion, or by itself protect against changes between reading and writing. Run the cases and describe how the real system would revalidate version or precondition at change time. Finish the exercise with the expected functional result and record required by the next shift. Express with best-effort logs or out-of-order status events does not by itself provide a complete ordered decision ledger. Choose persistence and reconciliation that meet the operation’s evidence requirement, and state which uncertainty still prevents closure.

# Original local decision exercise, not an AWS remediation engine or atomic write guard.
def plan_change(expected_id, current, desired, authorized):
 if not authorized:
 return "hold: authorization required"
 if current is None:
 return "hold: current state unknown"
 if current["id"]!= expected_id:
 return "hold: resource identity mismatch"
 if current["value"] == desired:
 return "no change: retain current-state evidence"
 return "prepare change: revalidate version before writing"

state = {"id": "route-a", "value": "healthy-path", "version": 7}
assert plan_change("route-a", state, "healthy-path", True).startswith("no change")
assert plan_change("route-a", state, "alternate-path", True).startswith("prepare change")
assert plan_change("route-a", None, "healthy-path", True) == "hold: current state unknown"
assert plan_change("route-b", state, "healthy-path", True) == "hold: resource identity mismatch"
assert plan_change("route-a", state, "alternate-path", False) == "hold: authorization required"
# Another actor can change version 7 after this read. Real writes need concurrency protection.
IN PRACTICE

Fictional example: an older trigger requests correction of an already compliant property. The runbook re-reads the resource, records evidence, and avoids an unnecessary restart.

Common pitfalls

Common mistakes: repeating without identifying effects, confusing HTTP 200 with job completion, losing input in Catch, assuming a new definition during redrive, and restarting based only on an old snapshot.

Related topics: Log investigation and evidence coverage

Take this idea with you

Recovery requires knowing what already happened, bounding what may happen again, and proving the resulting state.

Create account

Reference: DOP-C02 incident and event response objectives · DOP-C02

AWS is a trademark of Amazon.com, Inc. or its affiliates. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by AWS. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.