← Incident Management: coordinate and recover
03 / 6 · 40 MIN

Diagnosis and mitigation

Test hypotheses and reduce impact through proportionate, verifiable actions.

Concept and mechanism

Diagnosis requires comparing observations with expected system behavior. Build hypotheses that evidence can support or contradict. A release near onset is a clue rather than proof; a failure resembling last month's event does not guarantee the same cause. Record tests and results to avoid repeating paths already excluded. Choose checks that distinguish hypotheses with limited risk. A query may be preferable to a change if it provides the same information. When investigation requires a change, state the owner, expected state, stop condition, and recovery method. Runbooks must match the current version, dependencies, and conditions rather than merely the alert name.

Guided application

Impact reduction need not wait for complete causal analysis. In a fictional example, limiting a nonessential feature may restore capacity for priority operations if degraded mode is authorized and its effect measured. A rollback may help, but requires checking data compatibility and dependent changes; restarting every node can destroy availability and evidence. Preserve relevant information where feasible and assess the cost of heavy diagnostic collection. After the action, compare actual and expected results, including user operations, queues, and data. If the hypothesis fails, update the record and change direction. A successfully completed command establishes execution rather than functional recovery or removal of the cause.

IN PRACTICE

Completed rollback with persistent errors requires revisiting the hypothesis and mitigation.

Common pitfalls

Correlation as cause; automatic restart; out-of-context runbook; execution as recovery.

Related topics: Triage and impact · Coordination and responsibilities · Communication and evidence

Take this idea with you

Connect each action to a hypothesis, expected effect, and verification.

Create account

Reference: Hypothesis-driven troubleshooting · Incident management practices 2026-09; scoped Google SRE, PagerDuty and Atlassian examples