Draw an operation’s path
A missing confirmation does not necessarily mean payment never happened. In a fictional service, an event arrives, a function creates the posting, and final notification fails. Repeating everything without knowing state can create a second posting. Draw transport acceptance, invocation, persistent effect, and confirmation separately. Define a business identity retained across attempts and a persistent mechanism protecting the effect, including concurrency. Asynchronous Lambda can repeat an event even after success. A new identifier for every attempt provides attempt traceability but does not establish operation uniqueness. Record both identities and explain their relationship during diagnosis.
Capture failures before and after the handler
A backlog can age until events expire without entering the handler. Absence of stack traces therefore does not prove every operation ran. Track age, discard, execution, and outcome. An on-failure destination can retain the invocation record with request and response; a DLQ stores the original event and error attributes. These are different contracts. Test the capture path too: DestinationDeliveryFailures means expected information may not have reached the destination. Check permissions, types, and limits before trusting the investigation queue. A failure in the diagnostic mechanism has its own impact and belongs in the incident report.
Contain and restore with a work inventory
Setting reserved concurrency to zero can be an authorized containment measure. For new asynchronous invocations with failure capture configured, work is routed to the DLQ or on-failure destination without retries. Restoring capacity does not mean that work automatically reappears in normal execution. Before the change, identify where events will go and who will recover them; afterward, reconcile inputs, effects, and outstanding work. The same care applies to capacity limits: increasing concurrency can pressure an already degraded database. Define stop criteria and compare backlog age with actual downstream capacity rather than relying only on queue depth.
Prepare replay with verifiable limits
Native EventBridge replay returns to the source bus. For isolated tests, design a controlled route and confirm other rules do not produce unexpected effects. Bound time and relevant rules; events are not guaranteed to return in archive arrival order. The consumer must validate business identity and version. replay-name identifies replay but is not an operation uniqueness key. The archive retains events according to its retention; finishing replay neither deletes them nor proves every consumer completed. In the settlement case, compare identifiers, expected amounts, and confirmed states, using fictional data for the exercise.
Recover batches without acknowledging unperformed work
Lambda SQS partial responses require ReportBatchItemFailures on the mapping and the correct format returned by the handler. A whole-invocation exception fails the entire batch. For FIFO, stop at the first failure and return failed and still-unprocessed identifiers. In the exercise, A finishes, B fails, and C/D remain unprocessed; recovery includes B/C/D. The local model makes that decision visible but is not a deployable AWS handler. It implements neither persistence, concurrency, nor the service’s JSON response. Use it to discuss which effects are confirmed and which identifiers need another attempt.
Accept recovery through outcome
Configuration requires function timeout not to exceed message visibility. AWS guidance reserves more margin: six times function timeout plus the batch window. With 20 and 5 seconds, that gives a recommended minimum of 125 seconds. Distinguish guidance from minimum validation. Partial responses also do not automatically reduce polling when failures occur; control downstream pressure separately. To close the incident, assemble the expected set, completed work, repeated work, and what remains pending or quarantined. Operational completion depends on this correspondence and observed stability rather than merely an empty queue or COMPLETED status.
# Original decision model, not an AWS handler or a production replay tool.
def fifo_review(items, first_failure):
if len(set(items))!= len(items):
raise ValueError("Exercise identifiers must be unique")
if first_failure is None:
return {"completed": items[:], "retry": []}
if first_failure not in items:
raise ValueError("Failure must belong to the batch")
boundary = items.index(first_failure)
return {"completed": items[:boundary], "retry": items[boundary:]}
assert fifo_review(["A", "B", "C", "D"], "B") == {
"completed": ["A"], "retry": ["B", "C", "D"]}
assert fifo_review(["A"], "A") == {"completed": [], "retry": ["A"]}
assert fifo_review(["A", "B"], None) == {"completed": ["A", "B"], "retry": []}
assert fifo_review([], None) == {"completed": [], "retry": []}
A completed, B failed, and C/D depend on FIFO order. Recovery retries B/C/D, retains evidence of A’s success, and reconciles any partial effect from B.
Common pitfalls
Changing identity on every attempt; confusing asynchronous DLQ with SQS recovery; processing FIFO after the first failure; closing an incident on COMPLETED alone.
Related topics: Reliable signals and observability gaps
Repeated delivery is recovery only when it preserves identity, required order, and already applied effects, with subsequent reconciliation.
Reference: Lambda asynchronous invocation errors · DOP-C02