← Production Support L3: investigate and recover
01 / 12 · 20 MIN

Triage and organize the response

Connect symptoms to business impact and separate coordination, communication, and execution.

Impact before automatic severity

Start with the affected service, failed operations, impacted users or processes, known onset, and business deadline. A technical alert alone does not define severity. A staging failure and a payment-cutoff failure can share symptoms while differing in impact. Apply the organization’s matrix and update classification as information changes.

One coordination function, several competencies

Define who coordinates response, who leads operational investigation, and who keeps stakeholders informed. One person may hold several roles in a small incident, but the roles still exist. Maintain a shared timeline of hypotheses, actions, and outcomes. Avoid uncoordinated simultaneous interventions that obscure cause and effect and increase risk.

Communicate facts and uncertainty

An update should state confirmed impact, known onset or window, ongoing mitigation, and the next update time. Distinguish confirmed cause from hypothesis. Do not invent a recovery time merely to fill a field. You can commit to a next update even when no defensible recovery estimate exists.

Workplace application

In a fictional 06:10 incident, distinguish the unavailable service from the component generating the alert. Record affected operations, cutoff, and business alternatives. Agree an update cadence with the owner and keep someone available to execute authorized decisions. An unanswered question should remain visible with an owner and next evaluation time.

IN PRACTICE

“Since 06:10 UTC, valuation processing is delayed for three funds. The team is investigating database waiting; the cause is unconfirmed. There is no evidence of data loss. Next update at 06:30 UTC.” This separates hypothesis, impact, and assurance.

Common pitfalls

Classifying solely by message; promising recovery without evidence.

Related topics: Investigate hypotheses and dependencies · Recover batch chains without falsifying success

Take this idea with you

Severity follows impact and context; coordination enables recovery through traceable decisions.

Create account

Reference: Incident response · DR Production Support L3 2026.4; Linux, JDK 25 HotSpot, OpenSSL 3.5 and Kubernetes examples require installed-version checks