← Incident Management: coordinate and recover
08 / 8 · 60 MIN

State, handover, and coordination exercise

Maintain traceable decisions and continuity across teams using a simulation guide with changing context.

Separate proposal, decision, execution, and outcome

The document model requires an owner, decision=approved, and a stop criterion to mark an action ready. A proposal fails, and removing any mandatory field makes the check fail. These results validate a structural rule defined for the exercise. They do not authenticate authority, assess decision quality, or execute mitigation. In the operational record, separate the proposed option, decision taken, executor, execution time, and observed effect. Approval can precede an ineffective or partly effective attempt. This distinction prevents reporting recovery from an approved ticket and lets another team reconstruct what actually happened. Keep unknown outcomes explicit instead of filling them from expected results.

Confirm the state that was accepted

The fixture has state at revision three and an acknowledgement of revision two. The local rule rejects that handover because acceptance does not cover current context. Confirmation of revision three passes; if mitigation changes to revision four before effective transfer, the earlier confirmation is again insufficient for the new context. This does not require revision numbers in every tool; it makes visible what the next team acknowledged. In a real handover, confirm impact, active actions, abandoned hypotheses, limits, the next checkpoint, and ownership. A populated field does not establish that the person understood, has access, and can act. The human confirmation remains part of operational continuity.

Distinguish elapsed deadline from wall time

The exercise defines escalation after ninety seconds without the required response. Synthetic monotonic readings go from 1000 to 1120, so 120 seconds elapsed and the deadline was reached. Wall time stepped backward and would show minus thirty seconds in that subtraction. Monotonic clocks measure duration differences within appropriate scope; their reference is not a UTC time to publish in the timeline. Retain wall timestamps with a time reference for cross-team communication and choose a duration mechanism appropriate to timer requirements. The script does not change the computer clock or test system suspension or synchronization across hosts. These boundaries matter when translating the model into tooling.

Check dependencies of response tools

In the fictional graph, console and chat depend on identity-a and network-a. The phone alternative depends on carrier-b. When identity-a fails and other nodes remain available, the model leaves only the phone accessible. Different applications do not imply independence of identity, network, or power. The graph helps choose a check but does not prove that the number is current, someone answers, or dependencies are complete. Assign an owner to confirm the alternative route and keep incident state in a medium accessible to participants. Rehearsing contacts and continuity before failure reduces the need to discover the whole chain during response. The lab tests graph assumptions rather than real communication services.

Run a tabletop with phased information

For a facilitated session, assign coordination, operations, communication, and observation roles. Use only fictional data and do not contact real services. In the first phase, provide backlog 600, arrivals forty, completions one hundred, and an eight-minute cutoff. Request an update with assumptions and a decision. In the second, reveal that completions fell to seventy at minute four. In the third, propose limited admission and ask the team to identify displaced demand. Then announce chat unavailability and finally a mitigation change between acknowledgement and handover. The facilitator should retain decisions and questions without immediately providing answers. These phases are an authored training guide; no human group execution is claimed.

Evaluate decisions and follow through on learning

In the tabletop review, check whether the team updated forecasts when assumptions changed, distinguished mitigation from recovery, communicated displaced impact, confirmed an alternative route, and transferred the latest change at handover. The aim is not to reward whoever spoke fastest. Link each gap to an improvement with an owner, deadline, and observable condition: for example, repeat a handover with an intervening update and demonstrate that the recipient identifies the changed action and new stop criterion. The nine automated groups verify calculations and local rules; they do not measure understanding, cooperation under pressure, or operational competence. Keep human review and supervised application as outstanding work.

python3 content/labs/incident-evidence/run.py
# approved record!= executed action!= observed recovery
# acknowledgement revision 2 does not cover current revision 3
# monotonic elapsed: 1120 - 1000 = 120 seconds
# Dependency graph availability is not a real fallback test.
IN PRACTICE

The next shift accepted revision three, but a later change introduced another mitigation. The team communicates the delta and confirms updated state before transferring ownership.

Common pitfalls

Marking recovery from approval, accepting an old revision, using wall time for a timer without considering adjustments, and assuming another product has independent dependencies.

Related topics: Incident Manager · Change Management · Technical Project Manager

Take this idea with you

Continuity requires current state, acknowledged ownership, and means to act. Models expose errors; coordination competence needs practice and human assessment.

Create account

Reference: Managing incidents · Incident management practices 2026-09; scoped Google SRE, PagerDuty and Atlassian examples