1. Separate detection and delivery
A finding can exist without reaching the support team. Draw the flow: source, rule matching, target invocation, ticket creation, and owner assignment. Each transition needs different evidence. During an out-of-hours change, a finding screenshot does not demonstrate notification or shift acceptance. Use identifiers and timestamps to follow a test occurrence through to its destination. This lesson concerns delivery to EventBridge rule targets. Other mechanisms, including Custom Event Bus subscribers, Pipes, and Scheduler, have their own contracts. Do not transfer defaults or restrictions between them merely because they share the EventBridge name. Record the concrete delivery mechanism in the runbook so diagnosis does not rely on the wrong product.
2. Interpret retries and recovery queues
A retry policy bounds delivery time and attempts. When an event is dropped after that policy is exhausted, recovering the target does not automatically recreate the event. A DLQ retains failures for later handling. Some errors, including missing permissions, can reach the DLQ without retries. Read ERROR_CODE, TARGET_ARN, RULE_ARN, and, where relevant, EXHAUSTED_RETRY_CONDITION. If age expired first, increasing attempts alone leaves the same time window. For a rule target DLQ, use a standard SQS queue in the same Region. When configuring through the API, provide the required resource policy. Distinguish original target failure from a second failure while sending to the DLQ itself.
3. Recover through observable criteria
Define who monitors backlog, who repairs authorization, and who confirms ticket creation. Also observe InvocationsFailedToBeSentToDLQ: configuring a queue does not prove failed events reached it. After fixing the cause, rehearse a new occurrence and prepare handling of retained events. Reconciling identifiers and avoiding duplicate tickets is an operational decision in this scenario, not an automatic service guarantee. The local exercise classifies fictional observations into repairing DLQ delivery, repairing target authorization, preparing replay, or confirming downstream processing. It sends no events and does not reproduce the AWS state machine. Agree with RUN what happens if replay fails and which evidence permits incident closure.
4. Give missing metrics a meaning
A custom counter that publishes only errors differs from a heartbeat expected continuously. Silence can be normal for the former and indicate emission failure for the latter. Define that meaning before choosing TreatMissingData. notBreaching treats missing points as good; breaching treats them as breaches; ignore retains state when evaluation depends on those points; missing permits INSUFFICIENT_DATA when all points are absent. The retrieved range can contain additional real points. If enough exist for EvaluationPeriods, substitution is not used. Test interrupted emission and recovery, not just threshold violation. These examples concern ordinary custom-metric alarms, not log alarms or exceptions for particular metric namespaces.
5. Distinguish Macie inventory, selection, and analysis
Macie automated discovery chooses representative objects; it is not proof of full inspection of all data. For a migration decision, first identify the set needing analysis, then observe selection, permissions, and results. An object skipped for missing access is not an object analyzed without sensitive data. The error can involve object policy or the encryption key. Within a job, sampling depth concerns object selection, not reading a percentage of bytes from every file. Exclusion conditions take precedence over inclusion. Keep those choices in the acceptance package so a zero-finding result is not presented beyond the scope that produced it. Recheck coverage after repairing permissions.
6. Close coverage gaps and hand over to RUN
In a recurring job, excluding existing objects leaves history outside initial analysis; subsequent runs consider new or changed objects. Increasing frequency does not itself repair that scope choice. The current sensitivity profile is not an immutable archive either: subsequently changed or deleted objects leave the described statistics. For historical audit, retain evidence needed for the specific question with a date and scope version. In management reporting, separate what was configured, what could execute, what it found, and what remains unanalyzed. RUN handover should include owners, symptoms of degraded coverage, and criteria for repeating tests. A scoped conclusion remains useful when declared gaps still exist.
# Original fictional delivery triage, not an EventBridge implementation.
# No credentials, network calls, or AWS changes.
def triage(row):
if row.get("dlq_send_failed"):
return "repair-dlq-delivery"
if row.get("error") == "NO_PERMISSIONS":
return "repair-target-authorization"
if row.get("retained"):
return "plan-controlled-replay"
return "check-downstream-evidence"
assert triage({"dlq_send_failed": True}) == "repair-dlq-delivery"
assert triage({"error": "NO_PERMISSIONS", "retained": True}) == "repair-target-authorization"
assert triage({"retained": True}) == "plan-controlled-replay"
assert triage({"finding_exists": True}) == "check-downstream-evidence"
assert triage({}) == "check-downstream-evidence"
print("five delivery-evidence cases passed; no event was sent")
A rule creates zero tickets because the target denies access; the team uses DLQ attributes to locate the failure and confirm recovery through support.
Common pitfalls
Findings as notification; retries as permission repair; silence as universal health; sampling as full analysis.
Related topics: Telemetry and investigation · Governance and evidence
Accept results within demonstrated scope and identify the concrete stage where evidence stops.
Reference: Retrying delivery to rule targets · SCS-C03