Concept and mechanism
An SLI measures a defined service characteristic; the SLO establishes its target and window. Specify eligible events, success criteria, and measurement origin. An HTTP 200 indicator can be useful for transport but does not automatically cover content date, completeness, or correctness. Combine external journey signals with internal telemetry to locate failure. Mean latency can hide a slow tail; percentiles describe distribution rather than a fixed error percentage. Segment operation, version, and period when that separation helps explain impact. A measurement without samples should show the coverage limitation rather than automatically turn absence into zero failures.
Guided application
In the exercise, 19,900 good requests among 20,000 eligible requests produce 99.5%. For a 99.9% target, the budget allows twenty bad events; the hundred observed exceed it. This calculation uses events in a closed period and should not be directly converted into downtime minutes. In another case, fast responses return yesterday’s positions: add business-date validation and investigate the data chain. An alert should lead to action proportionate to risk and available time without requiring the team to watch a dashboard permanently. If collection fails during maintenance, communicate the unknown interval and seek alternative evidence before concluding that the service improved or stopped.
HTTP 200 with stale data requires content validation; zero samples requires measurement validation.
Common pitfalls
Excluding bad requests from the denominator; mixing time and events; using only the mean; treating absence as success.
Related topics: The service, impact, and ownership · Incidents, communication, and shift change · Batch, files, and reconciliation
Always explain what the indicator covers and what it leaves unobserved.
Reference: Implementing SLOs · DR APS professional curriculum 2026-09; vendor-neutral operational guidance reviewed 2026-09-30