A bool comparison does not filter the healthy instance
The third experiment uses error ratio 0.01 and threshold 0.02. The normal comparison returns no series because the condition is false. With bool, it retains an element valued zero. An alert rule considers elements present in the result, so the deliberately incorrect version becomes active. This helps explain why a dashboard expression might not work directly as an alert rule. Before enabling, test a healthy service and check whether the result should be empty. The lab file includes this incorrect rule to demonstrate behavior; it should not be published as an operational configuration. Inspect result shape as well as its numeric value.
Keep alert identity stable
The queue has values 10, 11, 12, and 13 at its first evaluations. The stable rule retains service and severity; the incorrect rule adds observed with queue value. Every change to that label creates a different identity and prevents completion of the three-minute pending period. An annotation can display the current value without changing identity. In the lab, the stable rule reaches firing at t=3m while the dynamic one is pending with observed=13. Do not confuse this identity with Alertmanager notification grouping. The first issue belongs to rule evaluation; grouping messages by service at a later stage does not recover lost pending duration.
Rehearse gaps and recovery on the timeline
for requires continuity of the active instance. In the eighth fixture, a stale series interrupts pending at t=2m; the condition returns at t=3m and completes three minutes only at t=6m. In another fixture, the rule is already firing when the condition ends at t=4m. With keep_firing_for set to two minutes, it remains firing at t=5m and ends at t=6m. Retention can avoid overly quick resolutions but does not replace investigation of missing data. The shift needs to distinguish current condition, retained state, and delivered notification. Record evaluation interval and rehearsal assumptions rather than promising exact timing beyond what the design supports.
Fix the population and interpret the budget
The canary fixture contains 100 eligible queries with two errors and 900 successful health checks. Mixing everything produces 0.2%; the query SLI produces 2%. Measurement must follow the agreed definition rather than whichever version permits approval. To interpret relative consumption, divide the bad-event fraction by the permitted fraction. In a separate example, 2% against 0.5% tolerance gives burn rate 4. That does not automatically mean 4% of monthly budget consumed or four minutes remaining. Those conclusions require period, population, and additional assumptions. Use results for a decision with defined authority and preserve evidence about the affected cohort.
Retain useful dimensions for release investigation
A service aggregate can be sufficient for global status but cannot reconstruct a version removed during aggregation. Plan which data supports canary versus previous-version comparison within retention and cost limits. During investigation, use traces with temporal relationships: children lasting 100, 200, and 300 ms can overlap inside a 350 ms request. Summing children does not give end-to-end latency. Examine waits and the path to completion, accounting for instrumentation limits. Coincidence with the release can guide a hypothesis, but rollback decisions should follow known criteria, authority, and compatibility. Preserve enough evidence to investigate after mitigation rather than relying solely on an aggregate screenshot.
Accept the complete observation and response chain
The lab executes ten original groups in promtool 3.15.0, with 20 expression checks and 12 alert-state checks, plus validation of five rules. It does not start Prometheus, scrape production, or send notifications. Passing these tests establishes fixture behavior rather than operator delivery or a Dynatrace installation. Handover should include an owner, runbook, permissions, freshness metrics, and an authorized routing rehearsal to a test destination. To declare batch recovery, confirm useful output and its date as well as the exporter. These examples are fictional and do not represent internal bank procedures. Independent specialist review and operational integration testing remain necessary.
groups:
- name: training-example
rules:
- alert: QueueStable
expr: dr_queue_depth > 5
for: 3m
labels:
severity: training
annotations:
summary: "Queue remains high"
observed: "{{ $value }}"The stable-label rule reaches firing at t=3m. Adding observed with changing queue value keeps new identities pending; moving that value to an annotation restores the distinction.
Common pitfalls
bool as a filter; current value in a label; for as sample count; firing as delivery; promtool success as acceptance of the whole platform.
Related topics: Incident management · SLIs and SLOs · Canary and rollback
A useful rule needs the correct population, stable identity, rehearsed timing, and a response chain demonstrated within authorized scope.
Reference: Alerting rules and states · Observability 2026-09; selected OpenTelemetry, Prometheus and Dynatrace Classic concepts