← AWS CloudOps Engineer: operations and recovery
01 / 7 · 30 MIN

Signals, alarms, and missing telemetry

Confirm what the dashboard observes before concluding the service is healthy.

Concept and mechanism

A metric is useful only when we know what it measures, where it comes from, and which series is evaluated. Normal CPU does not demonstrate free space in /var. Guest operating-system metrics need suitable collection, such as CloudWatch agent configured with correct permissions and destinations. Namespace, name, and dimensions identify the series. Replacing an instance can leave an alarm pinned to the previous InstanceId rather than the current resource. Check recent datapoints and each dimension’s meaning before changing thresholds. A chart containing old data does not prove present conditions.

Guided application

Also define what absence means. For a periodic heartbeat, stopping publication matters; for a sparse error count, absence may be normal. Missing-data treatment must match the signal and be rehearsed. Composite alarms correlate states but do not directly support every metric-alarm action, including EC2 and Auto Scaling actions. An APS handover should describe the complete detection path: emitter, series, rule, action, recipient, and expected response. Testing only dashboard presentation leaves notification and intervention capability unproven.

IN PRACTICE

After replacing EC2, confirm the new instance’s /var utilization series and interrupt publication in a controlled rehearsal to assess detection.

Common pitfalls

Green without data treated as health; CPU treated as disk utilization; alarms pinned to old resources; actions unsupported by alarm type.

Related topics: Performance and operational evidence · Reliability, scaling, and demonstrated recovery

Take this idea with you

An alarm must observe the right resource and have known behavior when its signal disappears.

Create account

Reference: CloudWatch alarms · SOA-C03; exam guide 1.1