← AWS Developer Associate: applications and operations
08 / 9 · 60 MIN

Event contracts, checkpoints, and recovery

Distinguish acceptance, completion, and recovery when combining Lambda, streams, S3 notifications, and workflows.

Design states the business understands

Imagine a fund close submitted through an internal interface. The button can accept the request before reconciliation finishes. Define distinct states for accepted, processing, completed, and uncertain outcome; each needs evidence and an owner. An HTTP 202 response to asynchronous invocation only supports moving to accepted. Retain operation identity and provide a status query. In handover, explain how to distinguish a request still pending from one requiring intervention to meet the closing deadline.

Verify recovery too

A failure destination is another operational dependency. An invocation record can help reconstruct what happened, but delivery can fail too. In an authorized rehearsal, induce a known failure and confirm the expected record and alert. If the queue is empty, compare that observation with DestinationDeliveryFailures before concluding there are no incidents. Define who may inspect data and how to avoid sensitive information spreading through support channels. The runbook should include the procedure when evidence collection itself fails.

Separate asynchronous queueing from streams

Do not automatically transfer asynchronous-invocation configuration to an event source mapping. For a stream, the mapping reads a shard and delivers a batch to the function; its checkpoint influences work resumed. Record who reads, who acknowledges, and what state persists between attempts. For a fund movement, business identity should survive batch repetition and a changed technical invocation ID. A design meeting should explicitly walk through failures before and after an external call.

Work through a partially successful batch

In the exercise, 201 completes, 202 fails, 203 fails, and 204 completes. With partial responses enabled, the lowest failed sequence establishes the retry start: 202. The effect of 204 can be requested again despite completing. Design a durable record that recognizes the previous outcome. Returning only failed IDs does not promise single execution of the others. Rehearsing a stop at the first failure can simplify the sequence but does not remove other duplication sources. Document that difference for the support team.

Prevent report-index regressions

An index maintained through S3 notifications should not treat arrival order as change order. Retain the sequencer per key, normalize length for hexadecimal comparison, and use a conditional update to prevent consumer races. For 0A followed by 9, compare 0A with 09 and reject regression. Do not compare those values across different keys. Decode the key according to the event format before requesting the object. Record why an event was ignored to distinguish expected duplication from lost work.

Calculate the deadline before choosing retries

The model uses three one-second executions separated by waits of two and four seconds: nine seconds overall. MaxAttempts=2 allows two retries beyond the first execution. With an eight-second deadline, the plan does not fit even before networking or overhead. The model excludes jitter, SDK retries, and scheduling; measure those effects in the real system. Error names matter too: States.TaskFailed does not cover States.Timeout. Define expiration handling and uncertain business status without assuming that ending a wait cancels an external effect.

Prepare replay and monitor recovery

Before replay, confirm stored-event compatibility with the fixed version and retain operation identity. Start with a sample whose outcome can be reconciled. If 120 records arrive per second and 100 complete, backlog grows by 20 per second; increasing retention only buys time. To recover 6000 records with constant arrivals of 120 and capacity of 150, net recovery is 30 per second and the model takes 200 seconds. Stop extrapolating if failures, variable load, or one shard dominate the lag.

IN PRACTICE

In a fictional close, the checkpoint repeats already completed record 204. The team recognizes business identity and verifies the existing outcome before contacting the supplier again.

Common pitfalls

Confusing 202 with completion; applying the SQS contract to a stream; ordering different objects by sequencer; ignoring inner retries in deadline calculations.

Related topics: Idempotency and data contracts · Backlog monitoring

Take this idea with you

Credible recovery requires recognizing prior effects, retaining failed work, and showing that the diagnostic path works.

Create account

Reference: Lambda with Kinesis · DVA-C02; exam guide 2.1

AWS is a trademark of Amazon.com, Inc. or its affiliates. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by AWS. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.