← Linux administration for operations
09 / 12 · 45 MIN

Isolate resources and establish readiness on Linux

Connect cgroup limits, resource pressure, and startup semantics to service-recovery decisions.

The service exists, but is it ready?

Startup has several boundaries: process creation, application initialization, dependency availability, and acceptance of a real operation. Write down which boundary the consumer needs to observe. In a fictional API, the port may open before the instrument catalogue has loaded. Type=simple does not express that functional condition. Type=notify helps only if the application correctly implements notification. After= orders jobs but neither starts a dependency by itself nor tests a SQL query. Design activation, ordering, and functional verification as separate decisions, including an exercise in which the dependency fails.

Quota, weight, and available capacity

Use the operational question to select the control. CPUQuota defines a CPU-time ceiling relative to one CPU; CPUWeight influences sharing among siblings under contention. In this exercise, a batch has CPUQuota=150% on an eight-CPU host: the ceiling is equivalent to 1.5 CPUs, not twelve. That also does not guarantee it can use all that capacity. Before requesting hardware, observe effective configuration, throttling deltas during execution, and batch progress. An approved change should also measure effects on neighboring services and preserve a rollback option.

Memory: locate the constraint

The host and unit have different budgets. In an exercise with 12 GiB free on the host and a 2 GiB leaf memory.max, global memory does not remove the local limit. Also inspect ancestors: a shared parent can constrain the workload before it reaches its own ceiling. MemoryHigh can be associated with reclaim and delays; MemoryMax can lead to OOM when usage cannot be contained. Record event deltas, interval, and cgroup identity. An OOM establishes a condition but does not alone distinguish a leak, legitimate spike, or inadequate sizing.

Measure stalls and avoid false attribution

PSI adds a time perspective: it shows time during which tasks lose progress because of resource pressure. some and full describe different conditions; neither is a percentage of occupied memory. Compare the same window with latency, workload volume, and hierarchy counters. In memory.events, an increase at a parent may come from a descendant; memory.events.local helps separate scope. Do not add overlapping averages or use a cumulative counter as the number of failures in the last minute. For thread-creation failures, also compare TasksCurrent and TasksMax, which include tasks rather than only processes visible in a summary listing.

Recover without multiplying effects

Before retrying a batch, identify the last confirmed boundary. If an instruction was accepted externally and the local checkpoint failed, a restart can duplicate effects. Use operation identity, state lookup, and bounded retry according to the actual contract. Start-limit-hit describes startup protection; clearing that state does not repair the initial configuration. Preserve evidence, address the observed cause, and validate resumption. In an authorized exercise, produce a worksheet containing the hypothesis, observation command, expected result, evidence against the hypothesis, mitigation, authority, and acceptance criterion. No limit value in this module is a universal production recommendation.

systemctl show funds-api.service -p Type -p MainPID -p ControlGroup -p CPUQuotaPerSecUSec -p MemoryHigh -p MemoryMax -p TasksCurrent -p TasksMax
systemctl cat funds-api.service
journalctl -u funds-api.service --since "2026-10-01 06:00:00 UTC" --until "2026-10-01 06:15:00 UTC" --no-pager
cat /proc/pressure/memory
IN PRACTICE

Fictional example: the API stops responding, local oom_kill increases from 0 to 1, and the kernel identifies the cgroup limited to 2 GiB. Reducing authorized concurrency recovers service. Record mitigation and investigate the memory profile; do not declare a leak fixed or change other services’ limits.

Common pitfalls

Confusing weight with quota, free RAM with absence of limits, active with readiness, reset-failed with repair, or restart with exactly-once delivery.

Related topics: Interpret services and logs with systemd · Distinguish space, inodes, and mounts · Interpret CPU, memory, and I/O waiting · Maintenance and demonstrated recovery

Take this idea with you

Locate the constraint in the hierarchy, measure impact in the correct window, and establish recovery through the operation the user needs.

Create account

Reference: systemd.resource-control(5) · DR Linux 2026.4; networking and Bash manuals reviewed 2026-10-01; cgroup v2 and upstream systemd manuals reviewed 2026-10-01; RHEL 10 examples; Linux man-pages 6.19; OpenSSL 3.5