← L2 Support: diagnose, mitigate, and escalate
04 / 6 · 40 MIN

Operational mitigation and validation

Choose controlled actions and demonstrate service recovery.

Concept and mechanism

A mitigation reduces impact without necessarily eliminating its cause. Choose an action compatible with evidence, runbook, and authority, identifying prerequisites, scope, risks, stop criteria, and recovery. Avoid combining changes when their results cannot be attributed. Restarting a process may release resources but can lose evidence or interrupt work; the decision needs context. A timeout on an operation with side effects does not prove that it failed before producing a result. Before retrying, correlate identity and state, establish idempotency guarantees, and bound attempts. Technical restoration should be followed by validation of outstanding work.

Guided application

In an exercise, 120 tasks arrive per minute and 150 complete after mitigation. With 900 tasks pending and constant rates, no retries or other arrivals, draining takes 30 minutes: net reduction is 30 per minute. This calculation does not establish deadline compliance if load changes. Check errors, latency, data freshness, and consumer results. In PostgreSQL 18, an active session with a populated wait_event can be blocked on a wait; do not conclude that it continuously consumes CPU. Lock observations may support escalation to the database team without unauthorized session termination.

IN PRACTICE

900÷(150−120)=30 minutes under the stated assumptions.

Common pitfalls

Timeout as no side effect; empty queue as all results correct; multiple simultaneous changes.

Related topics: Triage based on impact · Useful hypotheses and evidence · Diagnosis by layer

Take this idea with you

Validate technical recovery and business outcome separately.

Create account

Reference: Timeouts retries and backoff with jitter · Operational support; PostgreSQL 18, OpenSSL 3.5 and BIND 9.20.29 examples; reviewed 2026-09-30