Separate association from a useful comparison
In the sample, every east request uses the new version and every west request uses the old version. The error difference is real in the dataset, but region and version vary together. Comparing versions under similar conditions or obtaining other evidence is needed to separate regional dependency from code change. Containment need not wait for a definitive cause. A controlled comparison should keep operation, period, dependencies and load as comparable as possible, with limited exposure and authorization. Record unavoidable differences and avoid selecting only favorable periods. If rollback and consumer scaling happen together, later improvement does not establish that one intervention alone was sufficient. The timeline should retain both for subsequent investigation.
Advancement requires scoped criteria
This workshop’s decision function is deliberately simple: at least 50 observations, no observed errors, age up to 60 seconds and three relevant passing probes. These numbers are fictional, not an SRE standard or an SLA. Three successes block expansion for insufficient sample; one error among 50 blocks it through the stop condition; an observation aged 75 seconds blocks it for age. Even with 80 recent successes, coverage is missing if the east archive remains unvalidated. The allowNextStage result describes only the guide’s stage. It does not provide statistical certainty about absence of future failures. For low volume, the team should agree windows and checks representative of service cadence and criticality.
Rollback includes state compatibility
Availability of an old image does not guarantee that the earlier application can read data already written by the new version. Before choosing rollback, examine compatibility of schema, formats, contracts and consumers. In a fictional scenario, a required field changes format and the earlier version cannot read it. Reverting only the binary can widen unavailability. Restoring a backup is not an automatic consequence either: it may remove later work and require reconciliation. Discuss a compatible option with development and data owners, such as a fix, containment or selective disablement, respecting change authority. Define the expected functional result and when to stop. This lesson analyzes the decision; the lab performs no application deployment, migration or restore.
Hand over temporary conditions and closure criteria
Service may work under an exception that expires. In the fictional case, west receives work until 18:00 and the night team starts at 17:30. Handover needs an accepted owner for the decision before expiry, archive validation and diverted-load observation. Archive produced does not equal archive consumed; identify the team confirming the later result. In the postmortem, turn diagnostic failure into demonstrable criteria: correct rates with unequal cohorts, preserved cardinality after joins, a distinction between no data and zero errors, and blocking for missing samples or probes. The guide can support an English workshop with APS, development and business roles. No observed human execution or independent specialist review has yet taken place.
python3 - <<'PY'
import json
from pathlib import Path
e=json.loads(Path('content/labs/incident-analysis/evidence.json').read_text)
for name in ['missingBusinessOperation','smallSample','canaryStop','staleObservation','scopedNextStage']:
print(name, e['checks'][name])
PYA fictional rule permits the next stage after fresh observations and relevant probes but does not accept enterprise recovery.
Common pitfalls
Using region as a version control, expanding with insufficient samples, assuming reversible rollback and forgetting temporary exceptions.
Related topics: Monitoring and observability · Mitigation and RUN handover
Mitigation needs an observed result, explicit limits and pending work transferred to owners who accept it.
Reference: Canarying Releases · Incident management practices 2026-09; scoped Google SRE, PagerDuty and Atlassian examples