← Incident Management: coordinate and recover
11 / 12 · 70 MIN

Cohort impact and data quality

Investigate failures hidden by aggregation, filters, missing samples and late-arriving evidence.

Start with the operation and denominator

A dashboard answers a question defined by its query, rather than every question about the incident. Before quoting a rate, identify the operation, measurement point, window and unit: attempt, accepted request, intent or consumed result. In this workshop’s original dataset, east has 100 requests and 40 errors; west has 900 requests and nine errors. The total is 49 of 1000, or 4.9%. Simply averaging 40% and 1% gives 20.5% because it gives equal weight to cohorts with different volumes. Retain both the correct total and cohort breakdown. If east serves the cut-off report, a small proportion of volume can carry a large operational consequence.

Filters and absence change interpretation

The p95 also depends on population. The script computes nearest rank over synthetic values: in east, successful-request p95 is 120 ms; including slow failures makes it 6000 ms. The calculations do not contradict each other, but using the first to declare every attempt fast would be incorrect. After diversion, east has no requests. The query returns a zero count and NULL percentage; it has not observed zero percent failures in a positive population. Record no sample and seek a representative check of the operation. A passing online probe does not establish that the closing archive is produced and consumed. Coverage must follow the consequence that prompted the response.

Enrichment does not create new requests

The workshop associates tags with requests: east receives two and west receives one. The join returns 1100 rows for 1000 requests; summing errors across those rows produces 89 instead of 49. The defect appears because one row no longer represents one request. COUNT(DISTINCT r.id) fixes the denominator but does not automatically fix a numerator that still sums repeated rows. Also count distinct failing identifiers or aggregate requests before enrichment. Retain a test with a known result to detect query regressions. In APS work, this inspection helps when logs, inventory, teams and tags are combined in a report. More context should improve interpretation without silently changing the measurement unit.

Preserve what was known at each decision

Workshop event B occurs at 120 seconds and reaches the collector at 420. At 300 seconds, the team knows A and C; at 600, it knows B also belongs to the first window. Keep occurrence, arrival and report times. Update the count through an explicit revision while preserving what was available for the earlier decision. For adjacent windows, an inclusive start and exclusive end place the event at 300 seconds only in the second window. Querying arrival time answers a different question. The script executes these queries in in-memory SQLite over invented data; it does not reproduce production clocks or prove real synchronization. Use the results to prepare investigation and identify timeline limits.

python3 content/labs/incident-analysis/run.py --output /tmp/dr-incident-analysis.json
# SQL used in the synthetic in-memory dataset:
# SELECT region, COUNT(*) n, SUM(1-ok) errors
# FROM requests WHERE phase='incident' GROUP BY region
IN PRACTICE

The fictional dataset has 49 errors among 1000 requests, but 40 of 100 east requests fail near the cut-off.

Common pitfalls

Averaging percentages, confusing absence with health, counting enriched rows as requests and deleting revisions.

Related topics: Monitoring and observability · Mitigation and RUN handover

Take this idea with you

A rate is useful only when its counting unit, excluded population and observation availability are known.

Create account

Reference: Monitoring · Incident management practices 2026-09; scoped Google SRE, PagerDuty and Atlassian examples