← Incident Management: coordinate and recover
07 / 8 · 60 MIN

Impact, backlog, and forecasts

Measure impact and pending work, update forecasts when conditions change, and communicate mitigation scope.

Preserve uncertainty about onset

The fixture explicitly establishes that impact began between 08:57 and 09:00 UTC. Recovery was validated at 09:35, so duration is between 35 and 38 minutes. Detection at 09:07 occurred seven to ten minutes after onset; acknowledgement at 09:10 took three minutes from detection. These are different measures. Do not turn one incident into an average or use MTTR without defining boundaries. The exercise supplies these bounds: two isolated probe results do not automatically guarantee all impact occurred between them. In a real report, state how onset is known, which operations were observed, and which gaps remain.

Update the forecast when the rate changes

Backlog starts at 600 jobs. Forty arrive per minute and one hundred distinct jobs finish per minute, giving net reduction of sixty and an initial ten-minute projection. After four minutes, 360 remain. If completions drop to seventy while arrivals stay unchanged, reduction becomes thirty: another twelve minutes are needed, or sixteen from the start. At minute eight, 240 still remain. The script reproduces these two phases with exact arithmetic. It is not a benchmark or a complete queueing model. Use the projection as a decision condition and measure it again when load, mix, or capacity changes. An old estimate does not become a technical commitment merely because it was communicated.

Assess impact displaced by admission control

In an alternative scenario from the start, admitting ten of forty jobs per minute leaves ninety net completions per minute if capacity of one hundred persists. The 600 jobs drain in 20/3 minutes, about 6.67. However, thirty jobs per minute were deferred or rejected. The queue chart improves because some demand no longer enters; it does not establish full service restoration. Define what happens to that demand, who tracks outcomes, and when normal admission may return. If the deadline requires 115 completions per minute and the destination was tested only to 110, adding workers does not prove additional capacity. The decision should expose limits and consequences.

Count the cost of additional attempts

Forty business requests with three attempts each offer 120 attempts per minute. At a destination supporting one hundred attempts per minute, offered load exceeds fixture capacity. Do not confuse attempts, valid completions, and distinct requests. Immediate retries during saturation can consume resources needed for recovery. A timeout also does not prove the destination failed to execute an operation. Assess limits, intervals, client behavior, and retry safety before changing policy. In a fictional financial-instruction flow, uncertain outcomes need reconciliation and stable identity appropriate to the contract. The exercise calculates amplification; it does not implement retries or recommend universal backoff values. Keep those implementation questions visible in the mitigation plan.

Check freshness of the green signal

At 09:20, the dashboard retains healthy from a 09:15 observation. Its age is three hundred seconds, exceeding the fictional sixty-second limit. Evidence does not confirm current state, but neither does it independently prove unavailability: service and observability can fail differently. Request a current observation and suitable functional validation. Do not interpret absence of new alarms as absence of impact when collection may have stopped. Record the time of the measured event and dashboard inspection to distinguish old data from current presentation. The sixty-second limit belongs to the exercise and is not an SLA of any bank or provider.

Prepare an update supporting a decision

Write a short update containing the affected operation, relevant backlog, observed rate, forecast assumption, current action, and next information checkpoint. At fixture minute four, the projection is another twelve minutes if seventy completions and forty arrivals persist; that path will not meet the minute-eight cutoff. This information lets the business assess alternatives. Do not confuse the next update time with a recovery promise. If you committed to an update at 09:25, you can fulfill that commitment while ETA remains undetermined. Use language appropriate to the audience and retain detailed technical evidence in the appropriately controlled record rather than copying customer data into general communication.

python3 content/labs/incident-evidence/run.py
# Initial backlog: 600; arrivals: 40/min; completions: 100/min
# Minute 4: 360 remain; completions fall to 70/min
# Minute 8: 240 remain; revised total projection: 16 minutes
# All values are synthetic assumptions, not measured production capacity.
IN PRACTICE

A fictional funds service resumes accepting instructions, but deliveries remain delayed. The team recalculates the forecast and distinguishes recovered reception, ongoing drainage, and cutoff risk.

Common pitfalls

Using an old forecast after capacity falls, hiding rejected demand, counting retries as useful work, and trusting a healthy label with an old observation.

Related topics: Monitoring and Observability · Production Support L3 · Problem Management

Take this idea with you

Mitigation can improve one metric while retaining impact elsewhere. Communicate outcomes by operation with visible population, freshness, and assumptions.

Create account

Reference: Handling overload · Incident management practices 2026-09; scoped Google SRE, PagerDuty and Atlassian examples