← ITIL 4 Foundation: service and production decisions
09 / 9 · 70 MIN

Measure outcomes and improve flow

Define interpretable metrics and assess improvements through time, quality, and consumer impact.

Write the definition before calculating

An indicator needs a population, success condition, window, data source, and exception handling. If the business wants timely and complete reconciliation, measuring only server availability answers a different question. Agree who owns the commitment and what evidence supports it. An SLA contains the service commitment; indicators and targets help track it. The following calculations are fictional rules supplied by the exercise. SRE guidance supports operational reasoning and is not presented as a mandatory formula or a new official ITIL objective.

Count the outcome the consumer receives

Across 20 days, one late file and a different incomplete file leave 18 days meeting both time and quality: 90%. Do not replace that with infrastructure’s 99.99% availability. Both indicators can be correct while measuring different things. Retain the technical indicator for diagnosis and show the service outcome. If regions have different volumes, aggregate counts when the definition is per request: 900/1000 and 9990/10000 yield 10890/11000, or 99%. Also display regions rather than hiding A’s 90%.

Overlapping intervals count once

The fictional agreement excludes 120 maintenance minutes from a 43200-minute window, leaving 43080 eligible. Outages during 20–50 and 40–70 share ten minutes, so their union occupies 50 rather than 60. Calculated availability is about 99.884%, below the 99.9% target. In another exercise, only 02:00–02:30 is excluded maintenance: a 02:10–02:45 outage leaves 15 eligible minutes. Starting during maintenance does not automatically extend the exclusion. Record the rule, clock, and interval boundaries.

Show unknowns and the distribution tail

If 100 of 1000 eligible outcomes were unobserved, coverage is 90%. Do not declare all good because errors are absent or all failed without evidence; show the gap and define agreed handling. Percentiles also have limits. For 19 requests at 100 ms and one at 1000 ms, nearest-rank defined as ceil(0.95×20) selects element 19: p95 is 100 ms. That does not erase the slow request. The percentile definition belongs to the fixture; real tools may use other methods. Retain count, window, and distribution.

Follow time through handoffs

The request waits 90 minutes, receives ten minutes of work, waits another 60, receives 20 minutes of work, and ten of validation. Sequential total is 190 minutes: 150 waiting and 40 active. Automating half the activity saves 20 minutes, leaving 170, about 10.5% total improvement. Before promising 50%, observe queues, hours, and handoffs between teams. An approval can have a necessary purpose yet wait because information or coverage is missing. Improve the flow with evidence rather than removing controls because of their duration.

Compare improvement with quality and risk

A pilot with lower median closure time can generate more reopens. Compare equivalent populations and a follow-up window sufficient to observe outcomes. If incorrect permissions rise from 1% to 6%, containing affected automation and reviewing exceptions may be necessary before expansion. Define acceptance criteria, an owner, and user feedback beforehand. If a supplier promises only a response within four hours, do not use the internal team’s 20 minutes as a recovery guarantee within 60. The outcome depends on the full flow.

Guided practice and service reporting

Recalculate the fixture’s four results: 90% good days, 99% aggregated good requests, 50 downtime minutes without double-counting, and 170 minutes after partial automation. Then state what each number does not prove. The global indicator does not validate every region; availability does not guarantee a correct file; shorter time does not prove correct authorization. A useful report connects definition, value, data quality, impact, and the next decision. Local models validate supplied arithmetic and assumptions rather than real contracts, incident causes, or production performance.

Synthetic agreed rules, not universal ITIL formulas
good_days: 18 / 20
regional_good_requests: 900 / 1000 and 9990 / 10000
window_minutes: 43200; excluded_maintenance: 120
outage_intervals_minutes: [20,50), [40,70)
flow: wait90 + work10 + wait60 + work20 + validation10
automation: halves work and validation, leaves waits unchanged
IN PRACTICE

A technical dashboard is green, but only 18 of 20 days allow timely and complete close. The service review addresses the consumer gap and dependencies needing improvement.

Common pitfalls

Adding overlapping outages; changing denominators after failures; hiding unknowns; treating p95 as a maximum; promising total reduction equal to work reduction.

Related topics: Incident management · Service levels and improvement

Take this idea with you

A useful metric retains its definition, population, and limitations, and helps decide how to improve the service.

Create account

Reference: Service Level Management practice overview · ITIL 4 Foundation; syllabus v4.2.0 (March 2025)

ITIL® is a registered trademark of the PeopleCert group. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by PeopleCert. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.