← Storage: capacity, performance, and recovery
11 / 12 · 60 MIN

Measure I/O and synchronization contracts

Measure a bounded local workload and interpret logical operations, bytes, latency, and fsync without conflating application and device.

Define a bounded observable workload

The runner writes synthetic four-MiB files in a temporary folder. Each file receives 256 logical operations of sixteen KiB using the same known payload. Three profiles run twice, totaling 24 logical MiB. One operation may require more than one write call if the interface accepts only part of a buffer; the code continues with remaining bytes and counts calls. It does not assume a call equals one physical device operation. The Python file object is opened without its own buffering, while operating-system caches remain active. There is no direct I/O, cache eviction, raw-device access, or kernel-counter collection. The objective is learning to instrument and interpret a small workload without attributing its numbers to a production volume or capacity guarantee.

Keep synchronization contracts visible

The end-only profile requests fsync after all 256 operations. The every-32 profile requests it eight times, after each complete group. The every-write profile makes 256 calls. All three write identical bytes but perform different work. Batch duration starts after opening the file and ends after the last planned synchronization; it excludes close and verification reading. In end-only, final fsync time is stored separately and is outside individual write samples. In other profiles, the operation sample triggering fsync includes that call. Compare these definitions before interpreting p95 or MiB/s. Successful fsync and a matching readback hash are local observations; they do not demonstrate survival through power loss, hardware-cache behavior, or directory recovery after a crash. Those properties require a separately defined and authorized experiment.

Calculate rates and percentiles with explicit boundaries

The perf_counter_ns clock supports monotonic interval calculation. Nanoseconds are the returned unit rather than a promise of precision or instantaneous completion. The report divides bytes and logical operations by total batch time. Verification can confirm that operation rate multiplied by logical size corresponds to byte rate within this same boundary. It also retains every duration to calculate percentiles using a stated rule: nearest rank selects position ceil(p × n) in the sorted sequence, counting from one. With 256 samples, p99 uses position 254. Such a short sample contains little information about rare events and does not turn p99 into a guarantee. Durations include scheduling and overhead on the observed path; they do not isolate physical disk latency.

Repeat without inventing a universal ranking

The two repeats run in a fixed order with active caches. That limits causal comparisons between profiles. The runner requires neither that one profile always be faster nor that timings repeat exactly. It checks complete content, operation count, fsync count, and interval consistency. Each execution retains its actual values. To study application performance, prepare representative load, sufficient duration, a comparable baseline, controlled order or conditions, and visibility into errors and concurrency. A trial ending with an incomplete write is a failure even if its calculated rate looks high. The complete code below uses only its temporary folder and removes files afterward. The next lesson's models are calculated separately and do not come from this host's telemetry. They illustrate assumptions rather than replacing observation.

#!/usr/bin/env python3
"""Original bounded local I/O measurement lab: six 4 MiB files and four synthetic planning models."""
import argparse,hashlib,json,math,os,platform,statistics,tempfile,time
from fractions import Fraction
from pathlib import Path

def nearest(values,p):return sorted(values)[max(0,math.ceil(p*len(values))-1)]
def main(output):
 checks=[];trials=[]
 def check(label,ok):assert ok,label;checks.append(label)
 block=bytes(range(256))*64;count=256;expected=hashlib.sha256(block*count).hexdigest
 clock=time.get_clock_info('perf_counter');check('measurement clock is monotonic',clock.monotonic)
 with tempfile.TemporaryDirectory(prefix='dr-storage-measurement-')as temporary:
 root=Path(temporary)
 for repeat in range(2):
 for name,interval in [('end-only',0),('every-32',32),('every-write',1)]:
 path=root/f'{repeat}-{name}.bin'write_ns=[];operation_ns=[];sync_ns=[];written=0;calls=0
 with path.open('wb',buffering=0)as f:
 start=time.perf_counter_ns
 for i in range(count):
 op_start=time.perf_counter_ns;remaining=memoryview(block)
 while remaining:
 n=f.write(remaining)
 if n is None or n<=0:raise RuntimeError('No progress during bounded file write')
 written+=n;calls+=1;remaining=remaining[n:]
 write_ns.append(time.perf_counter_ns-op_start)
 if interval and (i+1)%interval==0:
 t=time.perf_counter_ns;os.fsync(f.fileno);sync_ns.append(time.perf_counter_ns-t)
 operation_ns.append(time.perf_counter_ns-op_start)
 final_sync=0
 if not interval or count%interval:
 t=time.perf_counter_ns;os.fsync(f.fileno);final_sync=time.perf_counter_ns-t;sync_ns.append(final_sync)
 elapsed=time.perf_counter_ns-start
 actual=hashlib.sha256(path.read_bytes).hexdigest;label=f'{repeat}-{name}'
 check(label+': complete byte count',written==count*len(block)==path.stat.st_size)
 check(label+': complete content hash',actual==expected)
 check(label+': positive elapsed and sample count',elapsed>0 and len(write_ns)==len(operation_ns)==count)
 check(label+': expected fsync calls',len(sync_ns)==(1 if not interval else count//interval))
 check(label+': operation covers its write time',all(a>=b>=0 for a,b in zip(operation_ns,write_ns)))
 check(label+': measured subintervals fit batch',sum(operation_ns)+final_sync<=elapsed)
 trials.append(dict(repeat=repeat,profile=name,blockBytes=len(block),logicalOperations=count,writeSystemCalls=calls,bytes=written,fsyncCalls=len(sync_ns),fsyncEvery=interval,elapsedNs=elapsed,writeNs=write_ns,operationNs=operation_ns,fsyncNs=sync_ns,finalSyncOutsideOperationNs=final_sync,operationP50Ns=nearest(operation_ns,.5),operationP95Ns=nearest(operation_ns,.95),operationP99Ns=nearest(operation_ns,.99),operationMeanNs=statistics.mean(operation_ns),logicalOpsPerSecond=count/(elapsed/1e9),MiBPerSecond=(written/1024/1024)/(elapsed/1e9),contentHash=actual,scope='Sequential local buffered OS I/O; Python file object unbuffered. One synchronous logical operation at a time. Batch timer includes fsync requests but excludes open, close and hash readback. Not device IOPS or power-loss durability.'))
 check('all trials use the same complete payload',len({t['contentHash']for t in trials})==1)
 check('temporary trial files removed',not root.exists)
 # These models are intentionally synthetic, not measured host/device counters.
 a=[1]*1000;b=[100]*10;combined=a+b
 percentiles=dict(unit='ms',samplesA=len(a),samplesB=len(b),p95A=nearest(a,.95),p95B=nearest(b,.95),meanOfP95=(nearest(a,.95)+nearest(b,.95))/2,combinedP95=nearest(combined,.95),combinedMean=str(Fraction(sum(combined),len(combined))),method='Nearest rank: sorted values at ceil(p*n), one-based. Original synthetic samples.')
 check('percentiles do not average into combined percentile',percentiles['meanOfP95']==50.5 and percentiles['combinedP95']==1 and percentiles['combinedMean']=='200/101')
 size=64;volume_iops=3000;volume_mib=200;host_mib=120;other_mib=40;available=host_mib-other_mib;limit=min(volume_iops,Fraction(volume_mib*1024,size),Fraction(available*1024,size))
 limits=dict(sizeKiB=size,volumeIops=volume_iops,volumeMiBPerSecond=volume_mib,hostMiBPerSecond=host_mib,otherMiBPerSecond=other_mib,remainingHostMiBPerSecond=available,theoreticalIops=str(limit),bindingBoundary='remaining host throughput',measured=False)
 check('host aggregate can bind before volume limits',limit==1280)
 free=180;burst=40;reserve=20;growth=30;lead=4;days=Fraction(free-burst-reserve,growth)
 capacity=dict(freeGiB=free,burstGiB=burst,reserveGiB=reserve,constantDailyGrowthGiB=growth,leadDays=lead,daysUntilReserve=str(days),remainingAtCompletionGiB=free-burst-growth*lead,extraTimeMarginDays=str(days-lead),measured=False)
 check('capacity plan has zero extra lead-time margin',days==4 and capacity['remainingAtCompletionGiB']==20 and capacity['extraTimeMarginDays']=='0')
 queue=dict(meanOutstanding=8,completedPerSecond=2000,meanResidenceSeconds=str(Fraction(8,2000)),assumptions='Stable comparable window, same population and boundary; uses means, not a percentile. Synthetic model, no host queue measured.')
 check('stable queue model gives four milliseconds',queue['meanResidenceSeconds']=='1/250')
 report=dict(python=platform.python_version,system=platform.system,release=platform.release,machine=platform.machine,runnerSha256=hashlib.sha256(Path(__file__).read_bytes).hexdigest,clock=dict(implementation=clock.implementation,monotonic=clock.monotonic,adjustable=clock.adjustable,resolution=clock.resolution),checks=checks,trials=trials,models=dict(percentiles=percentiles,limits=limits,capacity=capacity,queue=queue),totalLogicalBytes=sum(t['bytes']for t in trials),scope='Actual bounded sequential file timing on the recorded host; six 4 MiB trials, 24 MiB logical writes, plus four synthetic planning models. Timing varies with caches, scheduler, filesystem and host load. No direct I/O, cache eviction, raw device access, kernel I/O counters, Linux execution, network storage, power failure, production benchmark or universal performance ranking.')
 if output:Path(output).write_text(json.dumps(report,indent=2)+'\n')
 print(json.dumps(dict(trials=len(trials),checks=len(checks),models=len(report['models']),totalLogicalBytes=report['totalLogicalBytes'],output=output)))
if __name__=='__main__':
 p=argparse.ArgumentParser;p.add_argument('--output');a=p.parse_args;main(a.output)
IN PRACTICE

Tejo compares two reports using the same payload, but one synchronizes every write and the other only the batch. The team aligns the contract before declaring regression.

Common pitfalls

buffering=0 as absence of all caches; logical rate as physical IOPS; fastest run as guaranteed capacity; successful fsync as a power-failure experiment.

Related topics: Performance and measurement · Durability and acknowledgement

Take this idea with you

Compare work under the same contract and retain measurement scope. A rate is interpretable only when bytes, operations, and interval are defined.

Create account

Reference: Python operating-system synchronization interface · BigSavant Storage 2026-09; selected Linux and AWS storage behavior