Retries multiply work
A retry consumes capacity when a service may already be degraded. In an example with three layers and three total attempts per call at each layer, including the initial attempt, one operation can generate up to 27 calls to the deepest dependency. This is a persistent-failure bound, not a prediction for every request. Define where retries belong, which errors are recoverable, and a total time and attempt budget. Backoff spaces calls; jitter spreads clients that would otherwise wake together. Record effective library configuration rather than only runbook intent.
Timeout and unknown outcomes
A timeout can occur after the server accepted an operation but before its response arrived. For a fictional instruction, query state using stable identity and follow the idempotency contract. Repeating the same intent with a new key can create another effect. Reusing a key with changed data may also violate the contract. Confirm scope, key-retention period, and late-request behavior. Do not promise exactly once merely because an identifier field exists. Recovery must distinguish a confirmed result, confirmed rejection, and a state still requiring reconciliation.
Useful capacity and dependencies
Adding workers helps only if the dependency can complete more useful work. In a fictional incident, consumers increase from 20 to 60, but commits per minute fall and JDBC waiting rises. This is evidence against insufficient consumers, although the cause of waiting still needs investigation. Stop expansion under the plan and involve the dependency owner. Consider bounded admission or concurrency with understood impact, protecting critical work and keeping backlog visible. Compare valid completions per minute rather than only the number of active processes.
Rollout state is not rollback
In a Kubernetes Deployment, Available=True can coexist with Progressing=False and ProgressDeadlineExceeded. The first condition indicates minimum availability under the strategy; the second reports that rollout did not progress before the deadline. The controller does not automatically roll back because of that condition. A higher-level orchestrator may do so if configured, requiring separate evidence. Confirm cluster, namespace, revision, and events before preparing action. In an exercise, quotas prevent new Pods while the previous version keeps serving. Fixing capacity and rolling back are different decisions; neither is established by the warning alone.
Compare canary and reference versions
An aggregate can hide a failing version when it receives a small traffic share. Define metrics, comparable populations, and criteria before promoting a canary. In a fictional example, compare errors by operation, latency, and functional validation for new and reference versions over the same interval. Also establish whether the sample included the critical funds workflow; a few health checks do not represent it. If the gate fails, stop promotion and follow the authorized plan, considering data and configuration compatibility. A technically completed deployment does not replace service acceptance.
Close with reconciliation and RUN handover
Operational acceptance combines service, data, and support capability. In a recovery exercise, the application responds again but a valuation file still has unexplained differences. Keep reconciliation open, identify the expected population, and record approved exceptions. Give the next shift the active revision, temporary changes, pending replay requests, and closure criteria. The business owner needs cutoff impact without a promise based solely on technical metrics. Track recovery, reconciliation, and prevention as distinct states so RUN knows exactly what remains and who owns it.
In a fictional scenario, the canary version has 12% failures and the previous version 0.2%. The global average remains low because only 5% of traffic uses the new version. Analyze both populations before promotion.
Common pitfalls
Retry as a free operation; timeout as proof of no execution; Available as rollout completion; global dashboard as functional acceptance.
Related topics: Validate transfers and reconciliation · Measure impact and confirm recovery · Turn incidents into verifiable improvement
Recover capacity and outcomes while preserving operation identity and acceptance criteria.
Reference: Timeouts, retries, backoff and jitter · DR Production Support L3 2026.4; Linux, JDK 25 HotSpot, OpenSSL 3.5 and Kubernetes examples require installed-version checks