Choose the unit before counting
An HTTP attempt, a business request, and an incident are different units. The fixture contains eight attempts for four identified requests. Six attempts were unsuccessful; final outcomes are two successes, one failure, and one unknown. The 6/8 fraction describes attempts and cannot establish that 75% of requests failed. Group by business identity and define the evidence that determines final status. A client timeout can coexist with execution at the destination. In reporting, preserve the unknown: one failure among three resolved outcomes, with one of four still requiring reconciliation. Do not automatically turn that request into success or failure to complete the chart.
Normalize time without inventing accuracy
The pool warning has time 09:04 UTC; the client timeout has 10:05+01:00, equivalent to 09:05 UTC. Nominal order differs by one minute. However, each clock may be off by two minutes and possible intervals overlap. Converting offsets does not synchronize clocks. Keep original text, time reference, normalized instant, and known uncertainty. Then seek correlation IDs, call relationships, or local sequences that add evidence. Additional decimal places indicate display resolution, not guaranteed accuracy. Even an established precedence needs a plausible mechanism and additional observations to support a causal explanation, especially when several components are saturated. State uncertainty before using the timeline to assign corrective work.
Distinguish records from impact duration
Two records describe full unavailability of the same service during minute intervals 0–12 and 8–20. Adding gives 24, but minutes 8–12 appear twice. The union measures 20 minutes of continuous unavailability. Keep both records for causal investigation and measure impact under an explicit definition. This calculation assumes the same service, the same window, and full unavailability in both intervals. It does not permit indiscriminate addition of failures from different services or counting affected users without population data. If only one function was degraded, define the function and criterion before calculating. At handover, provide the event timeline and impact indicator with their respective units.
Compare cohorts when the mix changes
Before, the set has 10 failures in 1000 simple operations and 10 in 100 complex operations. After, it has zero in 100 simple operations and 90 in 1000 complex operations. Per-type rates fall from 1% to 0% and from 10% to 9%; the aggregate rises from about 1.82% to 8.18%. There is no contradiction: the more failure-prone type gained substantial weight. Actual aggregate impact must be reported. For an additional comparison, a hypothetical population with half of each type produces 5.5% before and 4.5% after. Identify these weights as chosen for analysis; they replace neither actual traffic nor isolation of the release causal effect.
See the population hidden by the percentile
The lab sorts 1010 latencies: 1000 values of 10 ms and ten of 1000 ms. Under empirical nearest rank, the p95 position is ceil(0.95×1010), or 960. That position contains 10 ms even though every specialized request took one second. A fast global dashboard does not invalidate that client complaint. Examine the population corresponding to the affected function and state observation count, window, and method. Averaging the two group p95 values gives 505 ms, not the combined p95; the mean of all latencies is not that percentile either. Different tools may use different conventions, so the method is part of the result.
Prepare reproducible evidence
The fixtures.json file contains fictional data; run.py computes eight groups and evidence.json retains results and hashes. Execute the script in a local Python environment and compare values before writing a conclusion. Do not add a statistical-significance claim: these numbers were constructed to demonstrate interpretation errors. In a real investigation, identify origin, window, filters, units, transformations, and limitations for each extract while following applicable access and retention rules. Keep original data separate from derived tables and link each assertion to the transformation producing it. A colleague should be able to reproduce the calculation and understand which conclusion remains open without relying on an isolated screenshot.
python3 content/labs/problem-evidence/run.py
# attempts: 8; requests: 4; final outcomes: 2 success, 1 failure, 1 unknown
# same-service unavailability: union 20 minutes, event sum 24
# nearest-rank global p95: 10 ms; specialized p95: 1000 msIn a fictional funds service, retries increase HTTP errors and complex traffic dominates the following period. Prepare two tables: per-request outcomes and per-type rates, retaining unknowns and volumes.
Common pitfalls
Counting attempts as requests, resolving unknowns for convenience, confusing UTC with accurate clocks, adding overlaps, and using a global percentile to dismiss an affected cohort.
Related topics: Monitoring and Observability · Incident Management · REST APIs
A useful conclusion retains unit, population, time, and uncertainty. Calculations can be correct while interpretation remains wrong if scope changes.
Reference: Effective troubleshooting · Problem management practices 2026-09; ServiceNow Brazil examples with scoped plugins and properties