Concept and mechanism
Daily volume, peak rate, and concurrency describe different dimensions. Dividing requests by seconds gives an average, not the peak the system must support. Define operation mix, data size, timing concentration, and dependencies. Average latency can hide a long tail affecting slower users. A percentile target requires measuring the corresponding distribution. More replicas add useful capacity only if work can be distributed and dependencies do not become the new limit. Connect sizing to quotas, connections, memory, storage, and scaling delay, including capacity during failures.
Guided application
In an original exercise, each instance was validated for 150 requests/s with the expected mix. Peak demand is 1200 requests/s. Under explicit assumptions of uniform distribution and linear capacity, eight active instances are required; nine must be provisioned to lose one and retain that capacity. This does not replace testing or additional margin. Load testing assesses expected volume; stress explores limits; spike testing assesses abrupt increases; soak testing finds degradation over time. Compare variants with equivalent data and traffic, measure errors and latency beyond CPU, and document simulated elements. A delay-free simulated dependency can hide the main bottleneck of the real path.
1200÷150=8 active instances; tolerating one loss requires 9 under the stated assumptions.
Common pitfalls
Daily average as peak; CPU as the only limit; assumed linear capacity; short tests as proof against memory leaks.
Related topics: Requirements and architecture decisions · Data and consistency · Caching, partitions, and queues
State assumptions and measure the path that must meet the target.
Reference: Architecture strategies for performance testing · System design patterns; PostgreSQL18 scoped examples; primary guidance consulted 2026-09-30