← Linux administration for operations
05 / 12 · 20 MIN

Interpret CPU, memory, and I/O waiting

Use metrics together to choose the next diagnostic hypothesis.

Load is not CPU percentage

Linux load average includes runnable tasks and certain tasks in uninterruptible wait, often associated with I/O. Interpret it relative to CPU count and workload. High load with moderate CPU utilization can indicate resource waiting; it does not automatically prove a CPU shortage. Compare samples over time rather than relying on one snapshot.

Free memory and usable memory

The kernel uses memory for cache; a small free value alone does not mean shortage. Observe available memory, swap activity, OOM events, pressure, and service latency. Distinguish virtual, resident, and shared memory when examining processes. Correlation with application behavior is more useful than selecting a universal threshold.

Samples, limits, and evidence

Tools such as vmstat, ps, and I/O metrics help separate hypotheses. The first vmstat line may summarize time since boot; subsequent lines represent requested intervals. A containerized process may be constrained by cgroups even when the host has resources. Before adding memory or killing processes, identify limits, concurrency, and what the action could interrupt.

Workplace application

For a batch exceeding its usual duration, align CPU, memory, I/O, and functional-latency samples over the same window. Record workload size to avoid comparing different loads. In containers, observe limits and events in the corresponding cgroup: available host memory does not invalidate a service reaching its own limit. Propose capacity changes after identifying the resource and expected effect.

uptime
free -h
vmstat 1 5
ps -eo pid,ppid,comm,stat,%cpu,%mem --sort=-%cpu
# Correlate with service latency, resource limits and the incident time window.
IN PRACTICE

The host has available memory, but a container is terminated for exceeding its memory limit. Lack of host-wide shortage does not remove the container limit. Investigate application usage and unit/container configuration before changing global capacity.

Common pitfalls

Reading load as CPU percentage, cache as a leak, or global memory as the container limit.

Related topics: Separate DNS, connections, TLS, and the application · Shell: arguments, pipelines, and exit codes

Take this idea with you

One metric suggests a hypothesis; combined metrics and context support it.

Create account

Reference: Linux cgroup v2 administration · DR Linux 2026.4; networking and Bash manuals reviewed 2026-10-01; cgroup v2 and upstream systemd manuals reviewed 2026-10-01; RHEL 10 examples; Linux man-pages 6.19; OpenSSL 3.5