Distinguish critical path from sum
Consider three calls with fixed durations of 40, 90, and 180 milliseconds, plus ten of work after receiving results. If they start together, do not contend for resources, and are all required, the model gives max(40,90,180)+10=190. If sequential, it gives 320. These numbers are neither measured percentiles nor inclusive of queues, connections, or concurrency limits. Draw dependencies before applying the formula: a call needing another call’s result cannot start simultaneously. In a fictional reporting service, comparison helps frame an experiment, but approval needs observed latency under representative load and criteria for partial results.
State the independence assumption
A different model asks for the probability that twenty dependencies finish within deadline when each has probability 0.99. Only under the independence assumption, with all required, does the result become 0.99 to the twentieth power, approximately 81.79%. The laboratory computes the exact fraction using Python; it measured no service. Dependencies sharing a database, network, or resource limit can have correlated outcomes. In that case, multiplying individual probabilities does not establish joint probability. Use the calculation to question expanding fan-out and request data. Also define whether an incomplete response is useful or whether the application must reject when a required result is missing.
Measure concentration by shard
The partition model has four shards, each supporting 300 requests per second. Distribution is 700, 100, 100, and 100: total demand of 1000 appears to fit aggregate capacity of 1200, but the first shard receives 400 more per second than it can serve. Spare capacity elsewhere does not automatically move to it. If the key is tenant_id and the dominant tenant remains in one partition, adding shards without changing distribution may preserve the problem. Request key-level and partition-level data and evaluate subdivision and cross-partition query costs. These calculations form a model; no sharding implementation was executed in this laboratory.
Size for cold-cache conditions
The fictional application receives 1000 requests per second. With a 95% hit rate, the model sends 50 to the origin. After cache restart, a 20% hit rate sends 800 to the origin, sixteen times the previous load. If origin budget is 200, calculated demand is four times that limit. Do not assume direct fallback preserves availability: it may overload the very dependency needed for recovery. Discuss admission, priorities, acceptable freshness, and degradation with the service owner. Then test startup and recovery conditions. The script only calculates values; it started no cache, measured no capacity, and validated no warm-up policy.
Preserve scope and authorization
Two organizational tenants can have a report with the same local identifier. A key containing only report_id loses that distinction and can return another tenant’s content. Review must identify every dimension changing the result, including tenant scope and, where applicable, authorized visibility. Separate keys do not replace authorization on access. Also assess data sensitivity and whether a shared cache is appropriate. When changing format, address old entries, invalidation, and compatibility between versions. This case is an architecture exercise using fictional data; it represents neither an observed incident nor an executed test against an actual cache.
Connect cost, outcome, and decision
Compare two fictional variants costing 600 in the same window and scope. The first receives 100000 requests and correctly completes 90000; the second receives 95000 and correctly completes 94000. Dividing by received requests supports a different reading from dividing by correct outcomes. If the objective is cost per correct outcome, use 600/90000 and 600/94000 while retaining quality, freshness, and deadline criteria. Finish with a decision record containing the assumption, evidence, gaps, owner, and review condition. The guide can support a meeting with architecture, APS, and FinOps, but that human session has not taken place and the figures are not actual vendor costs.
parallel_ms = max(40, 90, 180) + 10 # 190
sequential_ms = 40 + 90 + 180 + 10 # 320
# Probability below assumes independent outcomes:
from fractions import Fraction
all_within_deadline = Fraction(99, 100) ** 20
# Synthetic arithmetic, not a measured service percentile.
origin_warm = 1000 * (1 - 0.95) # 50/s
origin_cold = 1000 * (1 - 0.20) # 800/sA fictional architecture has 1200 requests/s aggregate capacity but concentrates 700 on a shard supporting 300. Total headroom hides a growing local queue.
Common pitfalls
Adding incompatible capacities, multiplying probabilities without independence, using average hit rate as a guarantee, or counting accepted requests as correct outcomes.
Related topics: Monitoring and Observability · Load Balancing · Technical Project Management
A calculation exposes an assumption to test. The decision needs representative measurements, business criteria, and explicit operating limits.
Reference: Performance testing · System design patterns; PostgreSQL18 scoped examples; primary guidance consulted 2026-09-30