← Incident Management: coordinate and recover
01 / 6 · 40 MIN

Triage and impact

Identify the affected service and decide when to mobilize coordinated response.

Concept and mechanism

Triage turns a signal into an impact description that supports action. A CPU alert, a user call, and a delayed batch are starting points rather than automatic severity classifications. Record expected behavior, observed behavior, known onset, affected operations, and the trend. Separate facts from hypotheses and identify what still needs confirmation. A small infrastructure failure can prevent a critical operation, while many alerts can describe the same cause without increasing affected customer count. Use the organization's agreed severity matrix. SEV levels used by PagerDuty or Atlassian are examples of their processes rather than universal rules or BNP Paribas policies.

Guided application

In a fictional example, only one queue is blocked, but it carries files needed for a cut-off. Do not assign priority solely from the percentage of healthy servers. State the deadline, dependencies, and consequence of not recovering. If several teams are needed or local response lacks capacity, mobilize coordination without waiting for a definitive cause. Classification can be revised as evidence emerges, retaining the decision record. Avoid spending the opening minutes debating a label while impact grows. When scope is uncertain, communicate that uncertainty and assign a check with an owner and deadline. Do not present an internal indicator as proof that every customer operation is working.

IN PRACTICE

One queue among dozens can block the only file needed for daily close.

Common pitfalls

Counting alerts as customers; vendor severity as a standard; waiting for cause before coordinating.

Related topics: Coordination and responsibilities · Diagnosis and mitigation · Communication and evidence

Take this idea with you

Prioritize service consequences using evidence and agreed criteria.

Create account

Reference: Severity definitions and coordinated response · Incident management practices 2026-09; scoped Google SRE, PagerDuty and Atlassian examples