← Production Support L3: investigate and recover
06 / 12 · 20 MIN

Turn incidents into verifiable improvement

Build a blameless analysis and actions with owners, deadlines, and effectiveness criteria.

Reconstruct what happened

Keep a factual timeline: detection, impact, decisions, actions, recovery, and validation. Separate observations from inferences and record gaps. Instead of stopping at “human error,” examine conditions that made the action possible or hard to detect: interface, permissions, review, documentation, workload, and available signals. The goal is improving the work system while retaining accountability for future actions.

Actions that change a condition

“Be more careful” does not define a verifiable change. A useful action identifies the condition to fix, owner, deadline, priority, and effectiveness evidence. Actions can target prevention, detection, containment, or recovery. Not all require automation: clarifying go/no-go criteria may matter. Prioritize risk effect rather than merely how easily tickets can be closed.

Automation with controls and handover

Before automating a runbook, specify prerequisites, inputs, permissions, expected outcome, errors, and rollback. Automation can amplify a wrong action if target or state is misidentified. Publish the revised procedure, communicate with shifts, and confirm teams can use it. Monitoring effectiveness in the next cycle or exercise differs from marking the task delivered.

Workplace application

Turn a vague recommendation into an observable condition. For example, automation should reject an environment inconsistent with the authorized request. The action includes implementation, an owner, and an exercise demonstrating rejection of the wrong target. Retain evidence for the valid case too, so the control does not block all work. Reviewing effectiveness differs from counting closed tickets.

IN PRACTICE

After a certificate failure, the action is to inventory endpoints and owners, alert at an agreed lead time, and rehearse renewal in a representative environment. Evidence includes the reconciled inventory and a test alert reaching the right team. “The incident was discussed” does not demonstrate prevention.

Common pitfalls

Measuring only document delivery; seeking blame instead of conditions.

Related topics: Shift handover and escalation · Runbooks that support decisions

Take this idea with you

A good analysis ends in verifiable changes and follow-through, not a comfortable explanation.

Create account

Reference: Postmortem culture · DR Production Support L3 2026.4; Linux, JDK 25 HotSpot, OpenSSL 3.5 and Kubernetes examples require installed-version checks