Choose the unit before calculating
Twenty deployments include two requiring immediate intervention and four unplanned incident-driven interventions. In this example, change failure rate is 2/20, or 10%, and deployment rework rate is 4/20, or 20%. Five associated alerts do not become five failed changes. Retain them for diagnosis and link them to the appropriate events. Recovery time from a failed deployment is also not automatically the duration of any telecom incident. Keep each metric’s population, period, and boundary explicit. If deployment appears before commit, investigate clocks, time zones, and event linkage before announcing improvement; replacing negative times with zero without explaining the rule creates misleading data. Measurement should support a service discussion rather than an individual ranking.
Track the complete journey
In a fictional service, the API answers correctly but users abandon reconciliation at the next step. Technical call success does not establish the journey outcome. Combine completion counts with investigation of difficulty. If there are 120 attempts, twenty retries, and one hundred unique eligible operations, ninety valid, completion per unique operation is 90%. Do not confuse attempts with new outcomes. Retain period-specific denominators: twenty complaints among one thousand users is 2%; ten among two hundred is 5%. Absolute count fell while the rate rose. Explain both facts before recommending action. Quantitative data locate a problem; observation and research can explain why people cannot finish the task.
Distinguish improvement from changed composition
Before a pilot, ninety of one hundred simple requests and forty of one hundred complex requests meet the deadline: 130/200, or 65%. Afterward, 810 of nine hundred simple requests and forty of one hundred complex requests do: 850/1000, or 85%. Neither type-specific rate improved. Only the share of the better-performing type increased. Present the aggregate and composition; do not attribute twenty percentage points to the method. Another pilot may introduce early validation and temporary APS support simultaneously. If time falls, the data do not isolate which change produced the effect. Record the hypothesis, choose a comparison acknowledging work differences, and identify further observations needed. A promising result justifies learning but does not remove the need to assess sustainability before expansion.
Define boundaries and transfer follow-up
The pilot cuts waiting from six to four days, removes six weekly development hours, and adds nine to APS. The service now consumes three extra hours. If it also exceeds the agreed escaped-error boundary, pause expansion and investigate the design. Before the experiment, define population, period, expected outcome, quality and workload boundaries, an owner, and the decision to continue, adapt, or stop. An elapsed deadline does not approve the practice. During project-to-service transition, adoption may only become observable at month-end. Identify partial coverage and transfer measurement to an owner accepting the date, evidence, and decision. Five error-free days do not establish behavior of an unexecuted cycle. End the lesson with a review note separating observations, inferences, and pending conditions.
review = {
'observed': 'waiting decreased; APS effort increased',
'not_established': 'effect without temporary support',
'next_decision': 'adapt pilot before expansion',
'pending': 'monthly reconciliation cycle',
}
# Synthetic case, not a measured organizational outcome.65% to 85% overall, with type-specific rates fixed at 90% and 40%: composition changed; within-group improvement was not demonstrated.
Common pitfalls
Alerts as deployments; retries as outcomes; aggregate as causality; speed without quality; unowned benefit after project closure.
Related topics: Capacity and forecasting · Dependencies and outcomes
An improvement decision needs comparable data, effects across the whole service, and explicit continuation criteria.
Reference: Measuring the success of your service · Service Manual current public guidance; Kanban Guide May2025; DORA five-metric model; primary references reviewed 2026-09-30