← AWS CloudOps Engineer: operations and recovery
02 / 7 · 30 MIN

Performance and operational evidence

Distinguish I/O limits, database waiting, and audit gaps.

Concept and mechanism

A growing I/O queue can have several causes. IOPS measures operations per second; throughput measures transferred data. For 64 KiB operations, 3000 IOPS represents 187.5 MiB/s, but a 125 MiB/s ceiling limits the result first. In practice, also check instance limits, latency, operation size, and application behavior. Increasing one parameter can leave another limit intact. For shared NFS files across Linux instances in several AZs, EFS Regional matches the requirement; a block volume or independent local disks do not automatically provide those semantics.

Guided application

In a database, low CPU does not exclude lock or I/O delays. Use Database Insights and available engine signals to relate load, waits, and SQL to the incident window. Confirm support and configuration before assuming a particular feature. In auditing, distinguish management from data events. A trail observing bucket changes may not collect GetObject. Data-event selectors should reflect approved needs, cost, and retention. Adding collection today does not retroactively recover events never stored. Communicate that limitation when investigation cannot reconstruct a past action.

IN PRACTICE

If batch time rose from ten to forty minutes while lock waits increased, identify competing transactions before concluding more vCPUs are needed.

Common pitfalls

IOPS treated as MiB/s; low CPU treated as no problem; storage selected only by capacity; new logging treated as recovering the past.

Related topics: Reliability, scaling, and demonstrated recovery · Stacks, images, and deployments

Take this idea with you

Choose intervention from the observed mechanism and document evidence limits.

Create account

Reference: EBS I/O characteristics · SOA-C03; exam guide 1.1