← Linux+: production administration and troubleshooting
08 / 8 · 70 MIN

Linux service pressure and limits

Interpret quotas, memory, and PSI before changing service resources.

Start with unit and interval

At 02:10, a batch is late while the global dashboard looks normal. Before proposing capacity, record which process performs the work, its unit, cgroup, and ancestors. Confirm cgroup v2, available controllers, and local versions. Keep two samples with a monotonic interval and instance identity. In this exercise, workers use fair-class scheduling and no burst. A path name may be reused after restart; that does not make counters from both instances comparable. The initial goal is a hypothesis that can be tested with a bounded change.

A budget shared by workers

CPUQuota=50% is a ceiling relative to one CPU, even on an eight-CPU host. With cpu.max=50000 100000, four workers accumulating 12,500 µs each consume 50,000 µs in this model. Parallelism can spend the budget sooner; it does not multiply it. Do not derive an exact pause or linear throughput improvement: scheduling, load, and other constraints affect the result. CPUWeight governs competition and AllowedCPUs limits where execution can occur. Neither name should be treated as synonymous with quota or a guaranteed reservation.

Inspect the constraining ancestor

A child set to max may share a parent slice with a finite quota. If two batches compete in that slice, changing only the child may not resolve the delay. Inspect siblings and parent; connect each counter to its documented scope. In the example, nr_periods increases by 100 and nr_throttled by 30. That means 30 additional throttled periods among 100 observed, rather than 30 lost seconds or 30% failed requests. Include latency and backlog in the handover to connect technical evidence with impact.

Memory: distinguish pressure, limit, and kill

A service can become very slow without losing processes. If MemoryHigh was lowered and high starts rising, investigate reclaim, pressure, and workload before diagnosing a leak. MemoryMax is a different containment mechanism: when usage cannot be reduced at the limit, group OOM can occur. Reading oom_kill alone does not identify origin: that counter includes group processes killed by different kinds of OOM killer. Correlate logs, victims, local limits, and ancestors. Free host RAM and an active state do not replace this analysis.

Do not attribute an aggregate event to the wrong unit

In the normal scenario without memory_localevents, a slice’s memory.events aggregates descendants; memory.events.local scopes local events. If the aggregate rises while the local value does not, investigate children. Record mount options because they can affect expected interpretation. A counter change indicates an event in the interval; it does not automatically provide the victim name or root cause. Avoid blaming the payments service merely because its name appears on a slice that also contains a reporting process.

PSI measures stalled time

Use memory.pressure to distinguish pressure from occupancy. some observes periods with at least some stalled task; full observes when all non-idle tasks in that scope are simultaneously stalled. Full is a subset of some: do not add them to get total pressure. If some total grows by 1,500,000 µs over 10,000,000 µs, the interval shows 15%. A full increment of 400,000 µs in the same window is 4%, included within that 15%. These calculations are neither RAM percentages nor latency percentiles. Compare them with requests, throughput, and load.

A bounded operational experiment

For the batch, compare a window under equivalent load while preserving counters and the initial result. The team may stagger jobs, reduce concurrency, or temporarily adjust quota at the authorized level. For memory, reverting a recently changed threshold may be more bounded than raising every ceiling. Define duration, owner, rollback conditions, and criteria beforehand: completion before cut-off, reconciliation, neighboring errors, and latency. Improving only the triggering metric is insufficient if work remains incomplete.

Guided exercise and shift handover

In the fixture below, calculate equivalent quota, the fraction of throttled periods, and both PSI fractions. Answers: 0.5 CPU, 30%, 15% some, and 4% full. Then consider the unit being recreated: establish a new baseline instead of turning a smaller counter into improvement. Hand over identity, interval, local and ancestor settings, hypothesis, authorized decision, and business outcome. Local calculations verify only these fictional inputs. Separate practice on an authorized Linux system is required to observe actual scheduling, memory, and distribution behavior.

Synthetic fixture; cgroup v2; fair-class; no burst
cpu.max: 50000 100000
instance: batch-A (unchanged)
interval_usec: 10000000
nr_periods: 1000 -> 1100
nr_throttled: 40 -> 70
memory.pressure some total: 2000000 -> 3500000
memory.pressure full total: 100000 -> 500000
IN PRACTICE

A shared-quota slice delays two batches despite idle CPU. Parent, siblings, and the business window are part of diagnosis.

Common pitfalls

Quota as reservation; weight as ceiling; high as a kill; aggregate events as local; adding some/full; computing deltas across instances.

Related topics: Resource diagnosis · Change management

Take this idea with you

Identify scope and interval, test a bounded hypothesis, and validate the service outcome.

Create account

Reference: Control Group v2 · XK0-006 V8

CompTIA® and Linux+ are trademarks or registered trademarks of CompTIA, Inc. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by CompTIA. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.