Concept and mechanism
An event-driven integration needs to distinguish delivery from business processing. EventBridge retries eligible failures with time and count limits. The default policy allows up to 24 hours and 185 attempts; it is not a promise of unlimited functional execution. Certain errors, such as missing permission, can go directly to a DLQ without retries. Read error metadata before assuming the counter is wrong. For rule DLQs, use SQS standard in the same Region and configure the policy permitting delivery. A queue created through the API does not guarantee that authorization is correct. Also observe failures to send to the DLQ itself.
Guided application
After correcting integration, accumulated work still needs recovery. Identify effects already applied, including through an alternative procedure, and use idempotency identifiers and reconciliation. Replay a bounded batch, observe results, and expand only with evidence. A DLQ retains messages; it does not eliminate duplicates or automatically replay all business work. During investigation, build a timeline of changes, traffic, metrics, and traces. Latency after a release is a hypothesis, not isolated proof of cause. Automatic remediation should have stop conditions: if repeated restarts worsen impact, contain execution and escalate with evidence.
Fixing permission resolves the delivery path. Retained movements need comparison with work already handled manually.
Common pitfalls
Retry as guarantee; DLQ as exactly once; technical fix as closure; correlation as root cause.
Related topics: Security and operational evidence · Pipeline, artifacts, and identity
Close the incident when service and pending work are reconciled.
Reference: EventBridge DLQ · DOP-C02