Aggregation needs distribution and counts
The synthetic model creates one thousand one-millisecond samples and ten one-hundred-millisecond samples. Group p95 values are one and one hundred. Their average is 50.5, but combined p95 is one millisecond under nearest rank: position 960 is among the first thousand samples. Even weighting percentiles does not reconstruct the required distribution. To aggregate, retain samples or compatible histograms with counts, declaring approximations and bucket boundaries. For the overall mean, use summed durations over total count; this model yields 200/101 milliseconds. A low global percentile also does not eliminate the slow group's problem. Report relevant segments and volume so that a large fast population does not hide impact on a smaller critical workflow. The fixture's numbers are invented teaching data, not production measurements.
Find the layer constraining the workload
Another model uses constant 64-KiB requests. Its fictional volume permits 3000 IOPS and 200 MiB/s. The host permits 120 aggregate MiB/s, of which other consumers use 40. Under these assumptions, 80 MiB/s remain, equivalent to 1280 operations per second at 64 KiB. That ceiling is lower than the volume's individual limits. Increasing only volume IOPS does not remove the aggregate constraint. The example does not represent commercial specifications for an instance or EBS type. In an actual investigation, confirm operation size and mixture, per-layer limits, competing load, and observation window. The application may combine or split requests before they reach the device. Use compatible counters and test the chosen intervention against service criteria, including latency and errors as well as average rate.
Relate queue, rate, and latency without mixing metrics
For a stable system and comparable population, the model relates mean outstanding requests, mean completion rate, and mean time within the same boundary. Its example uses eight mean outstanding requests and two thousand completions per second, yielding four mean milliseconds. This is neither a p99 estimate nor an automatic definition of physical service time. If queue length comes from a device while rate counts application calls, caching and aggregation can make populations different. A window spanning counter reset also needs treatment. Before calculating, record source, unit, interval, and semantics for every field. More decimal places cannot correct incompatible scope. The runner measures no actual queue; it evaluates only the arithmetic relationship under explicit assumptions. Its result should be reported as a model, not host telemetry.
Connect capacity to intervention lead time
The capacity model has 180 GiB free, a planned 40-GiB burst, a 20-GiB reserve, and constant daily growth of 30 GiB. Four days remain until reserve is reached. If expansion also takes four days, extra margin is zero. Late approval or above-forecast growth can compromise reserve. Define triggers, ownership, and an authorized alternative instead of treating the forecast as certainty. In a thin-provisioned design, also track pool data and metadata; logical filesystem space does not answer every constraint. The workshop below brings together four models and local measurements. Its forty-minute instructional duration is neither a recovery deadline nor an SLA. Production acceptance requires its own workload, platform, contract, and owners, which have not been exercised here. Keep that pending evidence visible during handover.
40-MINUTE GUIDE
0–8: Save the previous lesson's complete code as run.py in an authorized local folder. Use Python 3.13; the executed reference was 3.13.1 on Darwin 27.0.0 arm64. The experiment writes 24 logical MiB across six temporary files and makes local fsync calls.
python3 run.py --output./storage-measurement-evidence.json
8–20: Confirm six trials and 43 checks. Compare bytes, operations, fsyncCalls, and timing boundaries. Check the relationship between logicalOpsPerSecond, blockBytes, and MiBPerSecond. Inspect operationNs and finalSyncOutsideOperationNs. Require neither identical timings nor a fixed speed ranking.
20–30: Solve four models: combined p95 of 1 ms; a 1280 operations/s ceiling; four days until reserve with no margin beyond lead time; four mean milliseconds in the queue model. Explain assumptions and distinguish means from percentiles.
30–40: Write representative-trial criteria: workload, contract, baseline, window, errors, tail latency, concurrency, reserves, and acceptance authority. Identify caches, fixed order, and short samples as limitations. Confirm temporary-area cleanup. The workshop has not yet been performed with participants.Douro wants higher volume IOPS, but the host is already constrained by aggregate throughput. Ria has four days of usable headroom and an expansion taking those same four days.
Common pitfalls
Averaging percentiles into a global percentile; device queue combined with application rate; theoretical ceiling as guaranteed rate; execution lead time counted as extra margin.
Related topics: Capacity and growth · Migration, FinOps, and RUN
The decision depends on the right distribution and layer as well as lead time and assumptions. Keep models, measurements, and acceptance distinct.
Reference: Original I/O timing and planning fixtures · BigSavant Storage 2026-09; selected Linux and AWS storage behavior