Concept and mechanism
An alert should identify a condition requiring action and provide decision context. Define service, window, population, and outcome before choosing a threshold. For a 99.9% SLO, the permitted bad-event fraction is 0.1%. If the observed fraction is 0.5%, the burn rate is five: budget is consumed at five times the sustainable rate in that interpretation. This calculation alone does not say when remaining budget runs out; you need the period, accumulated consumption, and subsequent behavior. Combined long and short windows help detect meaningful consumption and recognize recovery. Documentation examples are not universal thresholds for services with different traffic patterns and consequences.
Guided application
In Prometheus, for keeps an alert pending until its expression remains active for the defined interval. It does not mean waiting that long after recovery. Add context and a runbook through annotations. Alertmanager can group notifications, inhibit alerts when another is active, and silence matching notifications for an interval. These features reduce noise; they do not repair the fault. In a fictional maintenance case, an overly broad silence can hide an unrelated incident. Review matchers, expiry, and ownership. With low traffic, one failure among a few requests yields a large percentage: interpret volume and impact together. For a daily batch, elapsed time since last success can be more actionable than an alert based only on HTTP traffic.
0.5% errors against 0.1% permitted corresponds to burn rate 5.
Common pitfalls
Universal threshold; silence as resolution; for as delay after recovery; percentage without volume.
Related topics: Signals and service outcomes · Metrics and queries · Traces and context
Notify with impact, context, and a clear action.
Reference: SLO burn-rate alerting and traffic limitations · Observability 2026-09; selected OpenTelemetry, Prometheus and Dynatrace Classic concepts