Concept and mechanism
An SLI describes a measurement and an SLO defines its target within a window and population. Choose indicators representing the service used by the business: accepting and completing an order can matter more than a process being alive. Define success, failure, justified exclusions, and data origin before analyzing results. Across one million eligible requests, a 99.9% target permits one thousand failures. After eight hundred, two hundred remain within that final total. The error budget informs priorities and change decisions; it is not an automatic contractual penalty. An SLA has its own scope and consequences. Avoid changing the population after an incident merely to improve the reported number.
Guided application
In a fictional investigation, normal CPU does not exclude a functional failure. Compare completion rate, latency, and errors against the release and dependencies. Metrics reveal patterns, logs reconstruct events, and traces with propagated context locate waits across services. Correlate request, version, and time without unnecessarily exposing sensitive data. Temporal coincidence guides a hypothesis but needs additional evidence. Alerts should reach an owner with a possible action and enough context to assess impact. To reduce noise, connect urgency to the service and the rate of objective consumption rather than notifying every fluctuation. Preserve the evidence used to confirm or reject hypotheses.
A request can wait on a dependency for 900 ms while CPU remains almost idle.
Common pitfalls
CPU as experience; correlation as cause; longer retention as context; SLO as automatic contract.
Related topics: Requirements, costs, and platform selection · Data, resilience, and events · Networking, provisioning, and capacity
Measure the outcome and follow the request to locate the wait or failure.
Reference: Service level objectives · Current linked standard guide; edition date unconfirmed (2026-09-30 inspection)