← Incident Manager: coordination, recovery, and learning
01 / 6 · 40 MIN

Declaration, impact, and priority

Turn scattered signals into a response proportionate to service impact.

Concept and mechanism

An Incident Manager organizes response when degradation requires coordination across people, teams, or decisions. The title does not mean being every component’s specialist or accepting any risk on the organization’s behalf. Define beforehand who may declare an incident, how responders are called, and which impact criteria guide priority. Alert count alone does not measure impact: a hundred alerts from one component can represent a single failure, while one missing file can threaten a critical delivery. Record what is known, what remains unconfirmed, and how impact is developing.

Guided application

In a fictional example, a funds batch fails at 16:10 and the agreed delivery is at 17:00. The server is available, but the expected result was not produced. Confirm affected services, dependent users or processes, time remaining, and recovery alternatives. If coordination across APS, database, and business teams is needed, activate the applicable process without waiting for a complete cause. Severity should follow the local matrix rather than numbers copied from a vendor. Reassess it with new evidence and explain changes. The practical priority is reducing impact while preserving an organized response and a factual basis for later decisions.

IN PRACTICE

Fifty minutes remain before the deadline, but that does not automatically mean every recovery option has fifty usable minutes.

Common pitfalls

Prioritizing by alert count; waiting for cause; copying external SEV levels; confusing server availability with completed delivery.

Related topics: Command, delegation, and shared state · Communication and uncertainty · Mitigation decisions and evidence

Take this idea with you

Classify by actual and expected impact using local criteria and explicit uncertainty.

Create account

Reference: Severity Levels · Google SRE incident guidance; PagerDuty contextual incident model; NIST SP 800-61 Rev. 3 April 2025; editorial review 2026-10-01