Start with the operational decision
In a fictional settlement service, a green line can mean requests without failures, no requests, or no telemetry. Before choosing a threshold, write down the decision the signal should support: call APS, interrupt promotion, or investigate lost visibility. Identify who produces data, when it arrives, and which population is measured. A count of log deliveries is not necessarily a count of unique business operations. Correlate identifiers without exposing sensitive data and retain enough evidence to explain a decision during shift handover. The initial question is what was observed and what remains unknown; dashboard color comes afterward.
Design the metric before the chart
A metric filter processes new events; creating one after an incident does not reconstruct points from the preceding hour. Retained logs remain useful for historical investigation. In a dimensionless filter, defaultValue=0 can represent a period with ingestion and no matches; without ingestion, that point is absent. With dimensions, that defaultValue is unavailable. Choose dimensions supporting useful decisions, such as service or environment, with controlled cardinality. A different requestId for every request would fragment series; keep that detail searchable in logs. Document unit, population, window, and absence semantics. A receiving operations team can then distinguish incomplete configuration from expected behavior.
Calculate rates with the right population
If A fails 2 of 10 requests and B fails 18 of 990, the combined rate is 20/1000, or 2%. A simple mean of percentages would give equal weight to unequal volumes. The local exercise calculates the combined rate and returns None when there are no requests: this keeps observed zero distinct from an undefined rate. In CloudWatch metric math, division by zero drops the point; it does not confirm success. Before using the result for an SLA, confirm that numerators and denominators measure the same operation, interval, and scope. A transport retry must not silently inflate the denominator of unique operations. The code teaches arithmetic rather than replacing business reconciliation.
Evaluate delay and missing data
A late metric can leave the latest period permanently filled with an artificial value. FILL with zero in an alarm requiring every point above threshold can keep it OK. Measure delay and exercise failure patterns; M-of-N with M below N may be appropriate, but requires a decision about tolerance and detection. Know the window too: sliding evaluation can use additional older real points when recent data has gaps. A wall-clock window uses its aligned interval without querying those older points. Record configuration and actually evaluated data before concluding that treatMissingData was incorrectly ignored.
Suppress notifications without hiding state
During maintenance, suppression can prevent a composite alarm’s actions while it remains ALARM. The team must know that silence was intentional and when it ends. WaitPeriod gives the suppressor time to follow the initial transition; ExtensionPeriod allows recovery after the suppressor leaves ALARM. Set these timers from observed behavior while retaining an owner and a way to identify conditions persisting after the window. Do not declare service recovery merely because notifications stopped. During operational acceptance, demonstrate notification when required, controlled suppression, and return to eligibility. Record who receives the signal and who owns the action.
Recover evidence continuity
Subscriptions can deliver repeated logs; a consumer producing effects needs idempotency. A failure such as AccessDenied can temporarily disable the filter and leave logs skipped during that interval. Correcting permission demonstrates that the path works again but does not prove the destination has complete history. In the practical case, the source retains the missing fifteen minutes. Bound the interval, recover events with preserved identity, and compare coverage, duplicates, and expected outcome. If complete recovery is impossible, declare the gap and impact. The RUN summary should identify observed data, evidence limits, tested alarms, and escalation conditions.
# Original local exercise: aggregate observations, no cloud calls.
def error_rate(samples):
for errors, requests in samples:
if not (isinstance(errors, int) and isinstance(requests, int)):
raise ValueError("Counts must be integers")
if errors < 0 or requests < errors:
raise ValueError("Invalid population")
total = sum(requests for _, requests in samples)
return None if total == 0 else 100 * sum(e for e, _ in samples) / total
assert error_rate([(2, 10), (18, 990)]) == 2
assert error_rate([(0, 0)]) is None
assert error_rate([(0, 100)]) == 0
try:
error_rate([(3, 2)])
raise AssertionError("Invalid count accepted")
except ValueError:
pass
APS finds an OK alarm despite 20 failures in 1000 requests. Review separates the 2% rate, metric delay, and FILL behavior before deciding acceptance.
Common pitfalls
Averaging percentages with unequal volumes; treating silence as zero; declaring history complete after fixing permissions; forgetting suppression affects actions.
Related topics: Event recovery and business effects
A reliable operational signal has explicit population, timing, and absence semantics. Evidence recovery requires reconciliation.
Reference: CloudWatch Logs metric filters · DOP-C02