← KCNA: Kubernetes and cloud native foundations
07 / 8 · 60 MIN

Workshop: capacity, eligibility and startup

Interpret stage-specific events and resources and build a verifiable operational hypothesis.

Start at the failing stage

During production handover, separate admission, scheduling, image retrieval, initialization and execution. A quota error can occur before a Pod is created; FailedScheduling calls for destination and resource evidence; ImagePullBackOff points to image, registry and credentials. A failing init may prevent application startup. Record the event and object before changing anything. In the fictional funds-check case, the business API has no logs because its process never started. Useful evidence is in the pull event. This sequence helps route the incident to the relevant owner and preserve the change window.

Capacity on the same destination

Prepare a per-node table of allocatable capacity and accounted requests. For a 500m/512Mi Pod, 800m/300Mi headroom lacks memory; 400m/1024Mi lacks CPU. Adding the two creates capacity offered by neither destination. This calculation is only a filter: labels, taints, volumes, ports and other conditions can also prevent placement. Use values from the same observation and state the exercise exclusions. In the APS meeting, distinguish overall capacity shortage from unsuitable distribution: the decision may involve moving workloads, measuring before revising requests or adding compatible capacity.

Resources across startup

Two application containers can run together, so add their requests. A regular init runs before them: in this example without sidecars, overhead or Pod-level resource settings, compare the largest init request with the application sum. A 1000m init and 300m/200m applications yield an effective 1000m, neither 1500m nor 500m. Do not generalize this simplification to native sidecars or other fields without reviewing their rules. After startup, a request is not a usage ceiling: on a node with resources, consumption can exceed it, subject to the applicable limit mechanisms.

Eligibility and maintenance

nodeSelector combines required labels. If it asks for disk=ssd and zone=west, a node with only disk=ssd does not match. A toleration permits a matching taint to be tolerated, but neither forces that destination nor provides memory. Cordon prevents ordinary new placements while leaving existing Pods; DaemonSets have specific behavior that needs attention. In a maintenance plan, identify applications still depending on the node and the authorized evacuation process. Do not declare the node empty merely because cordon completed. Acceptance needs observation of the relevant workloads.

Boundaries inside the Pod

Containers in one Pod share network space. Two processes attempting the same address and port can collide even if declared port names differ. Separating listeners requires process configuration or another Pod design. A regular init also defines a timing boundary: the application waits for completion. Investigate the correct container and stage. A Forbidden response when retrieving logs introduces another boundary: authorization for pods/log, separate from reading the Pod object. Changing a Service to public exposure does not automatically resolve any of these causes.

Hand over a verifiable hypothesis

Finish with object, symptom, hypothesis, evidence and expected post-fix outcome. For the reconciliation batch, record resource requests, an eligible node, approved image, completed init and functional result. Distinguish local placement arithmetic from cluster observation. This workshop model applies only stated filters; it does not run Kubernetes. The commands below are an inspection plan for a previously authorized context and namespace. An international team should be able to follow the reasoning from the record without vague statements such as “Kubernetes problem”. Include the next action owner and escalation condition.

kubectl get pods -n funds
kubectl describe pod funds-check -n funds
kubectl logs funds-check -n funds -c prepare-schema
IN PRACTICE

Node A:800m/300Mi; B:400m/1024Mi. A 500m/512Mi Pod fits neither despite aggregate headroom.

Common pitfalls

Adding node capacity, ignoring init, treating toleration as reservation or cordon as evacuation.

Related topics: Workshop: diagnosis, delivery and indicators

Take this idea with you

Locate the stage and confirm all requirements on the same destination.

Create account

Reference: Pod resource management · KCNA current four-domain curriculum; edition date unconfirmed

Kubernetes® and KCNA are trademarks or registered trademarks of The Linux Foundation. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by The Linux Foundation. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.