Start with the service and scarce resource
A fictional reconciliation service receives files, validates entries and writes results. CPU, memory, temporary disk, database connections and paid third-party calls are different resources. Identify what each request can consume and the consequence of exhaustion. A CPU quota does not automatically limit file size or the number of supplier calls. In architecture review, draw synchronous paths and queues, define who shares each resource and identify limit units. An aggregate ceiling can protect the service while still allowing one portfolio to consume all available capacity. The objective is useful work and predictable rejection under pressure, with coherent limits across layers.
Distinguish scheduling requests from runtime limits
In the inspected Kubernetes profile, requests inform Pod placement. A container may use more than its request when capacity exists and its limit permits. CPU and memory do not fail in the same way: CPU limits can cause throttling; memory limits are enforced reactively and can result in OOM termination. Do not classify both as simple slowness. A request is also not a business-latency promise. For rising memory usage, relate restarts, consumption, configuration and events before increasing the ceiling. A larger limit may fix undersizing or merely postpone a memory leak. Acceptance should include normal operation, peaks and behavior when a limit is reached.
Calculate using allocatable and all containers
The simplified worksheet has one node with 6000m allocatable CPU and 5400m already requested. Each new Pod has two ordinary running containers: an application requesting 300m and a helper requesting 100m. There are no init containers, overhead, other constraints or memory limitation in this calculation. 600m remain, each Pod requests 400m and only one fits by this dimension. Counting only the application would give two and be wrong. Even the result of one is only a CPU-based upper bound; in a real platform, affinity, taints, memory and other requirements also determine placement. Advertised physical capacity does not replace allocatable, which accounts for resources unavailable to Pods. The worksheet neither runs the scheduler nor measures performance.
Quota creates neither capacity nor complete isolation
In the same example, the namespace requests.cpu quota is 8000m with 6000m accounted usage. The remaining budget allows five 400m Pods, but the considered node has room for only one. Quota admission and scheduler placement answer different questions. Adding nodes also does not automatically increase namespace quota. An object rejected by quota is different from an admitted Pod Pending for lack of resources. Read the message, resource and failure stage before changing limits. A quota can contain aggregate consumption but does not replace authorization, network policy or kernel isolation. The acceptance record should identify the tested boundary and guarantees that remain outside scope.
Hand over observable limits and controlled recovery
If each replica opens a 20-connection pool and the database supports 100 for this service, six replicas can request 120. Adding Pods can worsen the downstream bottleneck. Coordinate pools, concurrency, bounded queues, deadlines and retry policies with dependency capacity. Do not turn an unbounded queue into a promise to absorb any peak. For RUN, record metrics by portfolio and operation, rejections, wait time, utilization, throttling and restarts, with owners and permitted actions. Rollback must consider already started work rather than only replica count. Summary: authorization, admission and capacity are distinct decisions; each needs evidence. Connect this lesson with observability, change, resilience and cloud costs.
FICTIONAL WORKSHEET, NO CLUSTER EXECUTED
Allocatable CPU:6000m; existing requests:5400m
New Pod:application300m + helper100m =400m
Remainder600m; CPU maximum:floor(600/400)=1
Namespace quota8000m; used6000m; remainder2000m
Quota maximum:floor(2000/400)=5
Other constraints excluded from this calculation
Six replicas x20 connections =120; database budget100
More replicas do not automatically increase downstream capacity.CPU remainder: 600m; Pod: 300m+100m=400m; at most one fits by CPU. Quota remainder: 2000m, or five Pods, without creating capacity.
Common pitfalls
Counting only the application; using physical CPU instead of allocatable; quota as reservation; more replicas without a connection budget; request as SLA.
Related topics: Capacity, availability and isolation · Authorization, replay and operational acceptance
Relate requests, limits and quotas to admitted work and actual service bottlenecks.
Reference: Resource Management for Pods and Containers · CCSP examination outline effective 2026-08-01; January2026 V2 PDF