← AWS Solutions Architect Associate: architecture decisions
17 / 18 · 60 MIN

Storage: IOPS, throughput, and persistence

Diagnose the storage path and connect units, concurrency, and persistence to the requirement.

Measure the business operation

A batch takes twenty extra minutes and an alarm shows low CPU. That does not establish a slow volume. Break elapsed time into data waiting, processing, remote requests, and writes. Collect I/O size and pattern, operation rate, latency, queue length, and volume and instance limits over the same interval. Distinguish a daily average from the critical closing period. Before changing resources, write a testable hypothesis and expected result. If the hypothesis is wrong, the next measurement should allow returning to diagnosis without accumulating paid capacity.

Connect operations and bytes

In a simple model, throughput in MiB/s equals IOPS multiplied by size in KiB and divided by 1024. Thus, 2000 operations of 64 KiB correspond to 125 MiB/s. The calculation assumes the stated effective size and excludes merging, splitting, and overhead. For a volume limited to 3000 IOPS and 125 MiB/s, 256 KiB operations reach the byte limit at 500 operations per second. This theoretical maximum is not a performance promise. Compare it with measurements and confirm that the application can issue enough work.

Find the shared instance limit

Two exercise volumes supply up to 200 MiB/s each, but the instance’s sustained EBS path supports 250 MiB/s in total. Demand of 350 leaves a minimum shortfall of 100. Buying more IOPS on one volume does not remove this limit. In production, inspect specifications for the actual instance and distinguish baseline from temporary capacity. Also check other volumes sharing the path. The conclusion should identify the limiting resource, observed workload, and analysis interval rather than simply choosing the largest value shown on a product page.

Match performance to access pattern

gp3 supports provisioning performance separately from storage capacity within supported limits. However, more configured capacity does not force a serialized client to generate requests. If each response triggers lengthy processing before the next I/O, investigate code and permitted concurrency. Do not add parallelism without checking ordering, locks, and downstream capacity. In a guided experiment, keep the volume constant and compare two client versions using the same dataset and correctness criteria. Measure duration, errors, and pressure on dependencies instead of seeking only a larger IOPS number.

Distinguish EFS choices

In EFS, performance mode and throughput mode answer different questions. General Purpose has lower per-operation latency; Max I/O is an older option that does not support Elastic throughput. Elastic adjusts throughput to activity without consuming burst credits, making it an option for irregular load. It still requires usage, cost, and regional-limit monitoring. In a production team, document both modes and the reason for each choice. A proposal to remove slowness should demonstrate that the problem lies in the file system rather than the client, network access, or an external dependency.

Connect persistence and recovery

A checkpoint helps only if it survives the considered event and can be used. Instance store loses data during stop/start, while reboot preserves data for that event. A regional EBS snapshot supports creating another volume in an AZ of the same Region; it does not start the application by itself. The exercise asks for batch recovery in another AZ: identify the snapshot and keys, create the volume, attach it to the correct environment, and validate functional state. This is a teaching plan, not an AWS execution. Also record how results after the recovery point will be reconciled before replay.

# Synthetic bound; no merging, overhead or cloud measurement
iops_limit = 3000
throughput_mib = 125
io_kib = 256
bound = min(iops_limit, throughput_mib * 1024 / io_kib) # 500
IN PRACTICE

Model: min(400, 250) MiB/s bounds demand of 350. The change must address the shared instance limit, not only one volume.

Common pitfalls

Confusing IOPS with MiB/s; treating peak capacity as sustained; ignoring the client; equating stop/start with reboot; declaring recovery complete merely because a snapshot exists.

Related topics: Data protection and recovery · Storage and content delivery · Costs, commitments, and retirement

Take this idea with you

Measure the complete path and retain units. Provisioned performance and functional recovery each need their own evidence.

Create account

Reference: EBS I/O characteristics · SAA-C03

AWS is a trademark of Amazon.com, Inc. or its affiliates. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by AWS. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.