Separate actors and permissions
The producer publishing to a topic, SNS delivering to a queue, and the consumer receiving and deleting messages perform different actions. To permit one specific topic, the queue policy can authorize sqs:SendMessage for SNS with aws:SourceArn matching that topic ARN. Do not confuse Publish authorization on the topic with SendMessage authorization on the queue. A simple consumer may need only ReceiveMessage, ChangeMessageVisibility, and DeleteMessage on the identified queue according to the described flow. Inventory additional implementation operations separately. Avoid using sqs:* to hide missing diagnosis. Also test denial for a source or action outside the requirement.
Include keys in recovery
Permission to read a queue does not establish permission to decrypt required data. In a DLQ containing messages encrypted at source, identify the keys protecting those messages and confirm kms:Decrypt for the consumer under applicable conditions. The DLQ’s current key should not be assumed to be the sole dependency of every older message. Consider identity and key policies together and retain appropriate least privilege. Do not delete a key merely because its source application retired: messages, copies, or other data may still depend on it. Operational review should connect retained-data inventory with keys, owners, and recovery conditions.
Diagnose filters and delivery failures
An SNS filter evaluates its configured scope. If it looks for kind in MessageBody but the producer sends it only as an attribute, absence from the body explains nonmatching. Policy changes can take up to 15 minutes to propagate; plan verification distinguishing propagation from a contract error. At another layer, EventBridge has time and attempt limits for target delivery. A target DLQ captures delivery failures when correctly authorized; it differs from a consumer DLQ after work is received. If attachment uses the API, confirm the resource policy allowing sqs:SendMessage for EventBridge and restricting the source rule.
Calculate net capacity
A queue absorbs temporary differences between arrivals and completions but creates no downstream capacity. In a constant model with 450 messages arriving per second and 600 completing, net reduction is 150 per second. A 9000-message backlog takes 60 seconds to disappear, assuming no retries or other constraints. Dividing by 600 gives 15 seconds and ignores new work. If each execution holds a connection and the budget is only 40 connections, proposing 100 concurrent executions violates that assumption. Measure useful throughput, work age, and downstream behavior. Quotas and actual capacity require separate verification; these numbers are synthetic rather than AWS measurements.
Reduce unnecessary work and requests
Long polling can reduce empty responses and return when messages are available, up to its configured limit. Also adjust the client HTTP timeout so it does not interrupt the intended wait. In an empty 60-second model, one call per second produces 60 requests; three complete twenty-second waits produce three. The 95% reduction belongs to that model, not a guaranteed bill forecast. Partial responses can also reduce repeated work: in a batch of ten with two failures, retrying only those two avoids eight already completed processing operations. Actual cost depends on calls, duration, size, traffic, and configuration. Compare alternatives with equal recovery guarantees and business outcomes.
Summary: validate the complete flow
Build a flow table: producer, topic or rule, delivery authorization, queue, consumer, key, effect, and acknowledgment. For each boundary, record observable failure, owner, and recovery. A received-message dashboard does not prove the final effect completed. An empty DLQ also does not establish absence of failures when delivery to that DLQ lacks authorization. In exercises, use known inputs to check filters, permissions, and retry sets, then compare expected and obtained outcomes. These local models teach decisions and allow arithmetic testing, but do not validate a real AWS environment. Keep assumptions explicit before turning a capacity or cost calculation into an operational decision.
Backlog of 9000, arrivals of 450/s, and completions of 600/s: net reduction of 150/s and sixty seconds in the constant model.
Common pitfalls
Confusing Publish with SendMessage, delivery DLQs with processing failures, permission configuration with testing, or concurrency with useful throughput.
Related topics: Events and decoupling · Identity and permissions
A design becomes operationally useful when delivery, authorization, capacity, cost, and recovery match the same expected outcome.
Reference: SAA-C03 performance objectives · SAA-C03