← Incident Manager: coordination, recovery, and learning
08 / 9 · 60 MIN

Decision and communication with incomplete evidence

Compare options, preserve uncertainty, and communicate what the recipient needs for a decision.

Compare hypotheses without losing mitigation

Network and application defend different explanations for an intermittent failure. Ask each workstream for the expected observation, an observation that would weaken its hypothesis, and a test distinguishing them. Comparing the same call on both sides may reduce uncertainty without changing the service. A complete cause is not required before assessing mitigation, but the decision needs expected effect, risk, dependencies, validation, and stop conditions. If restart and load reduction happen together, observed improvement does not isolate each contribution. Record progress and causal uncertainty separately. Before restoring load, define how relevant paths will be observed and when to revisit the option. Later investigation may explain the mechanism more deeply. Coordination seeks a defensible decision under current state without pretending a popular hypothesis became fact because it was repeated on the call. Ask what evidence would change the next decision, rather than requesting a longer technical narrative with no decision consequence.

Separate preparation, authorization, and effect

In the teaching scenario, A takes twenty minutes. B needs twelve minutes of isolated preparation and eight of activation. Activating both on the same target is incompatible. Preparing B while assessing A may preserve an alternative if resources and independence are confirmed; it does not authorize competing activation. Likewise, “diversion approved” does not yet mean “traffic diverted.” Request confirmation of execution and observed effect. If another service consumes the spare capacity supporting the decision, revisit feasibility before execution. An unknown outcome also needs attention: a reprocessing timeout does not prove that the destination had no effect. Preserve identity and check state with specialists before retrying. These distinctions help an Incident Manager ask concrete questions without replacing the people who know the technical mechanism. The record should show what was proposed, decided, started, completed, and validated. A timestamp on a decision is not a substitute for evidence of the resulting service state.

Produce an update that supports decisions

A business update needs affected capability, known impact, limits, ongoing work, and the next information point. A report received at 14:20 may describe 14:05 observations and not cover a 14:12 change. Preserve both times. If the next measurement is at 10:15, do not promise recovery at that time. When comparable rehearsals took eighteen to twenty-eight minutes, that interval can be communicated with assumptions and limits; the midpoint is not a guaranteed commitment. If business needs to activate an alternative, also explain when the decision is needed and what waiting implies. Public vendor references help structure communication, but their cadences do not automatically become local rules. Use language the recipient understands, avoiding a log list that requires them to infer impact. Technical specialists may need those logs through the appropriate route; business stakeholders need enough context to choose an operational response and understand what remains uncertain.

Correct expectations and handle contradictions

At 12:00 complete delivery was announced. At 12:07 reconciliation shows four missing files. Explicitly correct the assertion because business may already have acted on it. Do not wait for root cause before updating confirmed impact. One team may say “recovered” because the API responds, while another says “degraded” because reconciliation remains incomplete. Identify the capability and evidence behind each assertion before composing shared state. Vendor acceptance of investigation establishes neither a fix in progress nor a forecast. Write two English messages: a short business update and an instruction to the technical workstream. The first should explain impact and uncertainty; the second should request an observation, owner, and checkpoint. Use the example below as a practice structure and adapt it to scenario facts. This is a writing activity and sends no external message. Review whether a recipient could confuse your next-update time with a commitment to complete recovery.

Observed fact | Observation time | Limitation | Next evidence | Owner | Checkpoint
Decision: option, scope, assumptions, stop condition
Execution: started, completed, outcome validated or pending
IN PRACTICE

“Correction to our 12:00 update: four files remain unreconciled. The supplier has accepted the investigation; we do not yet have a supported recovery estimate. Our next update is at 12:20.”

Common pitfalls

Treating receipt as observation time; turning a checkpoint into an ETA; hiding correction; activating incompatible options; repeating an operation with unknown outcome.

Related topics: English communication · Hypotheses and evidence · Mitigation and validation

Take this idea with you

Communicate observed states and outstanding decisions while retaining conditions that may change the forecast or selected option.

Create account

Reference: Effective Troubleshooting · Google SRE incident guidance; PagerDuty contextual incident model; NIST SP 800-61 Rev. 3 April 2025; editorial review 2026-10-01