Compose timing and attempts
In the fictional path, three sequential stages reserve 150, 80, and 60 ms. They total 290 ms and leave ten before other overhead within a 300 ms limit. Using only the longest stage would fit a different model, not this sequence. Now consider three hundred-millisecond attempts with waits of twenty and forty: that is 360 ms, already beyond the limit. If the caller permits two total attempts and each triggers three SDK calls, six destination calls can occur. Always state whether a limit includes the initial attempt. AWS retry guidance supports examining load and synchronization; these numbers are original exercise data. Jitter spreads attempts over time but replaces neither limits nor guarantees success. Review policy already present in the SDK before adding another layer. The combined path, rather than each isolated configuration, determines the exposure that needs assessment.
Distinguish waiting, effects, and capacity
The client stops waiting at 300 ms, but the server records completion at 450. Timeout did not automatically cancel remote work. Define how deadlines propagate, what cancellation exists, and how to discover effects after response loss. Also distinguish a transient failure from a 403 confirmed as a permanently missing permission: repeating the same request with the same credentials does not address that condition. For asynchronous work, a queue can buffer a peak but does not increase processing rate. With two hundred arriving jobs per second, 150 departing, an initially empty queue, and twenty lossless seconds, one thousand jobs accumulate. Decision-making still needs peak duration, queue capacity, draining time, and treatment of delayed work. The arithmetic model neither measures actual throughput nor replaces a representative load exercise. Ask which assumptions change first when the dependency becomes slower under pressure.
Improve performance while retaining meaning and access
A cache can reduce latency while violating freshness or isolation. If the requirement permits data up to thirty seconds old and the design serves sixty-second-old data, present the conflict and alternatives to the requirement owner. If tenants A and B have report id=7, using only 7 as the key creates a collision. Adding tenant to identity avoids that particular collision but does not replace authorization on every access path. A check performed only during database queries can be bypassed by a cached response. Avoid another simplification in indicators: the average of two p95 values is generally not global p95. Prometheus documentation distinguishes already calculated percentiles from mergeable distributions. The model uses an explicit nearest-rank convention over complete observations; it simulates neither Prometheus interpolation nor actual metric collection. Keep performance evidence connected to correctness, freshness, and the access conditions required by the service.
Turn a local result into useful review
Use the final groups of the previous lesson's code to reproduce 360 ms, six calls, one thousand jobs, and the percentile difference. Each instance has twenty observations: A contains nineteen observations of 1 and one observation of 100; B contains twenty tens. Under nearest-rank, their p95 values are one and ten, their average is 5.5, and global p95 is ten. Ask a colleague to explain the difference using the distribution before changing code. Then write an English recommendation with hypothesis, available evidence, risk, and next exercise. Do not block on personal preference when two options meet the criteria; do not accept a 300 ms guarantee supported only by a local function either. Review should produce an understandable decision, owners, and information still needed. This discussion exercise is prepared but has not been performed by a human team. Local execution establishes model results, not service behavior, access enforcement, or the colleague's professional competence.
Hypothesis | Evidence | Version combination | Deadline | Expected failure | Next exercise | Owner
Questions: which time was omitted? which consumer is missing? which effect may already exist? which authorization applies?“One outer attempt can consume 360 ms before additional overhead. The proposed 300 ms commitment is unsupported. We need a combined retry policy and representative failure measurements.”
Common pitfalls
Waiting excluded from deadlines; layer attempts added; timeout treated as cancellation; queue as capacity; cache key as authorization; average percentiles as global percentile.
Related topics: Retries and backoff · Observability · Isolation and authorization
Review the complete composition and preserve result meaning and access boundaries when optimizing.
Reference: Timeouts, retries, and backoff with jitter · Google Engineering Practices, SRE and DORA; Microsoft architecture decision and collaboration guidance; OWASP threat modeling; UK lead developer framework; inspected 2026-10-01