Write the definition before calculating
An indicator needs a population, success condition, window, data source, and exception handling. If the business wants timely and complete reconciliation, measuring only server availability answers a different question. Agree who owns the commitment and what evidence supports it. An SLA contains the service commitment; indicators and targets help track it. The following calculations are fictional rules supplied by the exercise. SRE guidance supports operational reasoning and is not presented as a mandatory formula or a new official ITIL objective.
Count the outcome the consumer receives
Across 20 days, one late file and a different incomplete file leave 18 days meeting both time and quality: 90%. Do not replace that with infrastructure’s 99.99% availability. Both indicators can be correct while measuring different things. Retain the technical indicator for diagnosis and show the service outcome. If regions have different volumes, aggregate counts when the definition is per request: 900/1000 and 9990/10000 yield 10890/11000, or 99%. Also display regions rather than hiding A’s 90%.
Overlapping intervals count once
The fictional agreement excludes 120 maintenance minutes from a 43200-minute window, leaving 43080 eligible. Outages during 20–50 and 40–70 share ten minutes, so their union occupies 50 rather than 60. Calculated availability is about 99.884%, below the 99.9% target. In another exercise, only 02:00–02:30 is excluded maintenance: a 02:10–02:45 outage leaves 15 eligible minutes. Starting during maintenance does not automatically extend the exclusion. Record the rule, clock, and interval boundaries.
Show unknowns and the distribution tail
If 100 of 1000 eligible outcomes were unobserved, coverage is 90%. Do not declare all good because errors are absent or all failed without evidence; show the gap and define agreed handling. Percentiles also have limits. For 19 requests at 100 ms and one at 1000 ms, nearest-rank defined as ceil(0.95×20) selects element 19: p95 is 100 ms. That does not erase the slow request. The percentile definition belongs to the fixture; real tools may use other methods. Retain count, window, and distribution.
Follow time through handoffs
The request waits 90 minutes, receives ten minutes of work, waits another 60, receives 20 minutes of work, and ten of validation. Sequential total is 190 minutes: 150 waiting and 40 active. Automating half the activity saves 20 minutes, leaving 170, about 10.5% total improvement. Before promising 50%, observe queues, hours, and handoffs between teams. An approval can have a necessary purpose yet wait because information or coverage is missing. Improve the flow with evidence rather than removing controls because of their duration.
Compare improvement with quality and risk
A pilot with lower median closure time can generate more reopens. Compare equivalent populations and a follow-up window sufficient to observe outcomes. If incorrect permissions rise from 1% to 6%, containing affected automation and reviewing exceptions may be necessary before expansion. Define acceptance criteria, an owner, and user feedback beforehand. If a supplier promises only a response within four hours, do not use the internal team’s 20 minutes as a recovery guarantee within 60. The outcome depends on the full flow.
Guided practice and service reporting
Recalculate the fixture’s four results: 90% good days, 99% aggregated good requests, 50 downtime minutes without double-counting, and 170 minutes after partial automation. Then state what each number does not prove. The global indicator does not validate every region; availability does not guarantee a correct file; shorter time does not prove correct authorization. A useful report connects definition, value, data quality, impact, and the next decision. Local models validate supplied arithmetic and assumptions rather than real contracts, incident causes, or production performance.
Synthetic agreed rules, not universal ITIL formulas
good_days: 18 / 20
regional_good_requests: 900 / 1000 and 9990 / 10000
window_minutes: 43200; excluded_maintenance: 120
outage_intervals_minutes: [20,50), [40,70)
flow: wait90 + work10 + wait60 + work20 + validation10
automation: halves work and validation, leaves waits unchangedA technical dashboard is green, but only 18 of 20 days allow timely and complete close. The service review addresses the consumer gap and dependencies needing improvement.
Common pitfalls
Adding overlapping outages; changing denominators after failures; hiding unknowns; treating p95 as a maximum; promising total reduction equal to work reduction.
Related topics: Incident management · Service levels and improvement
A useful metric retains its definition, population, and limitations, and helps decide how to improve the service.
Reference: Service Level Management practice overview · ITIL 4 Foundation; syllabus v4.2.0 (March 2025)