Concept and mechanism
An SLI should represent the outcome relevant to service users. Low CPU does not guarantee fast requests; a slow dependency can keep processes waiting. Define eligible requests, success, boundaries, and window before calculating the SLO. For a 99.9% objective across two million requests, the allowance is two thousand failures. After fourteen hundred, six hundred remain. Burn rate compares observed error rate with the permitted rate: 0.8% divided by 0.1% is eight times. These calculations support decisions but require complete data and a stable definition. Do not change the denominator during an incident to hide unfavorable results.
Guided application
In a fictional funds operation, an alert combines sustained and recent consumption. Using AND between long and short windows helps confirm that the problem remains active, while a long window alone can remain elevated after recovery. Low-traffic services need interpretation of volume and complementary signals. During an incident, coordinate mitigation, evidence, and communication; complete root cause can follow limiting harm through a compatible action. During recovery, confirm the entire path, including queues and files, rather than only a successful restore. Finally, identify repetitive work growing with every batch: engineering a lasting correction can reduce toil. Scaling consumers also requires measuring dependencies; forty pools of twelve can compete for 480 connections against database headroom of 300.
A rollback can recover the service, but recovery ends only when relevant outcomes are valid again.
Common pitfalls
CPU as experience; missing data as zero; scaling without dependencies; all operations as toil.
Related topics: Organization, identity, and visibility · Infrastructure, revisions, and environments · Pipelines, promotion, and recovery
Measure the outcome and use the budget for explicit reliability decisions.
Reference: Alerting on SLOs · Current linked guide; edition date unconfirmed (2026-09-30 inspection)