← Incident Management: coordinate and recover
10 / 12 · 70 MIN

Partial recovery and reconciliation

Coordinate service resumption and prior-outcome handling without confusing mitigation with closure.

Separate new operations from history

Restoring the flag makes /ready and a new operation respond successfully. Both effects of the earlier intent remain present. This combination starts the second lesson: observed functional improvement exists alongside unfinished data work. A team can organize two workstreams with owners reporting to one coordinator. One observes stability and restart criteria; the other reconciles intents, outcomes and required actions. Do not delete a row merely because the count appears excessive. In an enterprise application, a record may correspond to effects in other systems and need an approved compensating action. The exercise asks learners to state what they know and who can decide what remains unknown.

Resume backlog under defined criteria

The fictional scenario contains 600 unacknowledged attempts, not necessarily 600 new intents. Some attempts may repeat the same work, and others may already have confirmed server outcomes. Group by stable identity, retain attempt identifiers and distinguish known from uncertain outcomes. Define resumption pace and a stop condition compatible with demonstrated capacity. An idempotency key does not remove network load, queries or database contention. If the cut-off cannot be met, communicate a conditional forecast and remaining impact. The decision should combine retry safety, capacity and ownership instead of treating one control as automatic authorization to release the entire queue.

Communicate and transfer each workstream

A useful update can state that new operations were validated at 14:20 while reconciliation remains assigned with a defined next checkpoint. An update time is a different commitment from a promise of complete recovery. In the workshop, write the message in English for an international coordinator, identifying scope, observed result, uncertainty and next step. Then prepare handover. If the next team accepted only traffic observation, do not assume it also accepted decisions about repeated effects. Record who received each workstream and the criteria for completing it. The document supports transfer, but human understanding and acceptance need observation in a facilitated exercise, which has not been performed here.

Turn learning into acceptance

The postmortem should not end with a vague action to improve retries. The material proposes observable criteria for the studied mechanism: lose the response after commit, repeat the same intent, introduce concurrency and send different content for an existing key. Check effects, returned results and solution limits. In a service calling an external system outside the SQLite transaction, add tests for that boundary; the local result does not automatically cover the new path. The workshop also asks learners to identify monitoring and communication gaps with owners and completion evidence. This course’s automated tests check the script and data. They do not measure team competence during a real incident or replace independent specialist review.

# Inspect retained results without starting another server.
python3 - <<'PY'
import json
from pathlib import Path
r=json.loads(Path("content/labs/incident-runtime/evidence.json").read_text)
for k in ["restoredTrafficDoesNotUndoPriorDuplicates",
 "reconciliationRetainsDifferentOutcomes"]:
 print(k, r["checks"][k])
PY
IN PRACTICE

A new operation passes after mitigation, but two earlier effects for the same intent remain in the database.

Common pitfalls

Deleting duplicates without a decision, removing replay limits because a key exists or transferring only monitoring while forgetting reconciliation.

Related topics: Diagnosis and mitigation · Communication and RUN handover

Take this idea with you

Functional recovery and reconciliation need their own criteria and owners linked to the same incident state.

Create account

Reference: Managing Incidents · Incident management practices 2026-09; scoped Google SRE, PagerDuty and Atlassian examples