Capacity in the failure state
In a fictional design, three zones each support 180 operations per second. Normal operation provides 540, but losing one zone leaves 360. If the requirement is to maintain 400 without provisioning resources, capacity is short by forty operations per second. Having a majority or several servers does not resolve that difference. Budgeting should consider capacity that remains usable in the defined failure mode, including shared dependencies. A plan that expects resource creation during the incident adds assumptions about time, quota, and control. Record those assumptions and the accepted degraded mode rather than presenting the normal total as recovery capacity.
Backlog competes with new arrivals
This lesson’s original model uses constant rates and uniform cost. With 4200 pending items, arrivals of 250 per second, and capacity of 320, seventy per second remain for backlog recovery. Calculated time is 4200 divided by seventy, or sixty seconds. Dividing by 320 would assume arrivals stopped. If arrivals equal capacity, positive backlog does not disappear; if arrivals exceed capacity, it grows. These results are model arithmetic rather than performance measurements. Variation in cost, latency, concurrency, and destination limits can change the behavior of a real application. Keep these assumptions attached to any reported estimate.
Additional attempts consume headroom
Add forty attempts per second with the same unit cost to the previous example. Headroom falls from seventy to thirty and idealized time rises to 140 seconds. An attempt may be necessary, but it consumes resources before producing useful results. Define eligible errors, elapsed-time limits, maximum attempts, and backoff with jitter in the real contract. A persistent authorization error needs corrected permissions rather than faster repetition. Also distinguish library retries from application retries. The lab only calculates the effect of an additional rate; it implements no client, jitter scheduler, or complete retry policy.
A smaller queue does not prove valid completion
An asynchronous API can acknowledge queue admission before the consumer produces an effect. Report accepted and completed as separate milestones. Observe item age, business deadline, class, errors, and exception handling. Removing items shrinks backlog without demonstrating correct execution. In a fictional funds flow, an expired instruction may require investigation or an authorized decision; the capacity formula does not determine whether it should be dropped, compensated, or still executed. The model also maintains no individual deadline-aware queue. The decision exercise must identify these additional requirements and assign responsibility for validating them.
Indicator population and timeline
The lab compares one hundred logical operations with 120 attempts. If 96 successful attempts correspond to 96 distinct operations, logical success is 96%, while per-attempt success is 80%. Both numbers can be useful, but they answer different questions. Define the population before an incident and retain identities for effect reconciliation. On the timeline, an example without overlap adds fifteen seconds for detection, twenty-five for decision, forty for reconnection, and twenty for validation: one hundred seconds to the final milestone. Do not report only the shortest stage or election while the consumer still cannot complete the agreed operation.
Workshop: a commitment before cut-off
Organize forty-five minutes with APS, business, and technical-manager roles. During the first fifteen, calculate residual capacity and drainage time using the supplied rates. Then introduce forty retries per second and a set of expired items. Use fifteen minutes to revise the estimate and decide which questions remain before promising recovery. During the final fifteen, communicate state, the conditional estimate, items without confirmed outcomes, and the next update in English. Acceptance requires representative exercises and functional reconciliation. Passing thirty-two local checks does not demonstrate actual load, zone loss, or independent review.
net_capacity = capacity - arrivals - extra_attempt_load
recovery_seconds = backlog / net_capacity # only when positive
# Constant rates and equal unit cost are exercise assumptions.
# No finite drainage time for positive backlog when net_capacity <= 0.4200 items, arrivals of 250/s, and capacity of 320/s: sixty seconds in the model. Another forty equal-cost attempts/s raises the estimate to 140 seconds.
Common pitfalls
Dividing by gross capacity; ignoring retries; counting attempts as operations; confusing accepted with completed; deleting items to improve a graph; presenting the model as a benchmark.
Related topics: Failure domains and residual capacity · Quorum and writer isolation · Replacement, identity, and acceptance
Recovery needs useful capacity and valid effects. An estimate must retain its assumptions, measured population, and consumer criteria.
Reference: Use static stability · BigSavant HA 2026-09; Pacemaker 3.0, etcd 3.6, PostgreSQL 18 and selected Kubernetes/AWS behavior