← Incident Management: coordinate and recover
05 / 6 · 40 MIN

Handover and recovery

Transfer responsibility and validate service before declaring recovery.

Concept and mechanism

A shift handover should transfer responsibility rather than merely send a document. The receiver needs to understand current impact, hypotheses, ongoing actions, risks, dependencies, and the next decision. Obtain explicit acknowledgement and announce the change to participants. In an international setting, use timestamps with timezones and clear language, avoiding phrases such as tomorrow morning without a reference. If the replacement has not acknowledged, keep responsibility visible and use the agreed escalation mechanism. Fatigue is also an operational risk: plan replacements before the team loses decision quality. Preserve knowledge of attempted actions so the next team does not repeat changes that worsened the situation.

Guided application

Recovery requires outcome evidence. An endpoint can respond while a backlog still prevents deadlines being met. Confirm representative operations, data freshness, integrity, and queue evolution using recent, healthy telemetry. In a fictional example, the service accepts files again but 400 deliveries remain pending; communicate partial recovery and track draining rather than declaring full normality. Acceptance criteria and observation periods should suit the service without importing a universal number. Nonurgent work can continue after the call ends if it has clear ownership and state. Distinguish mitigation, confirmed recovery, and permanent correction when updating the record and stakeholders. Explicitly retain residual risks that the next operational shift must monitor.

IN PRACTICE

Accepting new requests does not establish that pending deliveries completed.

Common pitfalls

Email as accepted handover; cleared alert as recovery; omitted backlog; unowned shift.

Related topics: Triage and impact · Coordination and responsibilities · Diagnosis and mitigation

Take this idea with you

Close each transfer and recovery with verifiable confirmation.

Create account

Reference: Incident command, state and explicit handoff · Incident management practices 2026-09; scoped Google SRE, PagerDuty and Atlassian examples