← CKA: Kubernetes administration and troubleshooting
06 / 7 · 24 MIN

Diagnose workloads with evidence

Interpret state, events, and logs before choosing mitigation.

Concept and mechanism

Pending, ImagePullBackOff, and CrashLoopBackOff describe different situations. Start with state, events, and last termination to locate the failure phase. For a restarted container, current logs may omit the earlier cause; kubectl logs with --previous and the correct container and namespace may retrieve it when available. ImagePullBackOff requires reading the image retrieval error: reference, credentials, network, and registry can explain different causes. A denied event makes pull authorization a priority investigation. Avoid changing probes for an application whose container has not even been fetched.

Guided application

OOMKilled directs analysis toward memory, limits, and usage without alone proving a leak or insufficient sizing. Correlate load and release and assess a mitigation’s node impact. Liveness may restart containers; readiness indicates traffic eligibility; a startup probe supports handling slow startup before activating the other checks. For scheduling, requests must fit an eligible node. Two nodes with 700m free each cannot host a 1000m-request Pod by splitting it between them. Retain evidence and stability criteria after restoring service.

IN PRACTICE

kubectl logs recon-7 -c app -n funds --previous queries the previous execution. Compare logs with kubectl describe pod recon-7 -n funds instead of repeatedly restarting without preserving the cause.

Common pitfalls

Confusing symptoms with cause; removing limits indiscriminately; expecting readiness to restart; summing node capacity for one Pod.

Related topics: Nodes and control plane · Cluster, context, and maintenance

Take this idea with you

Choose the next observation from the failure phase and recorded state.

Create account

Reference: Debug Pods · CKA Kubernetes v1.35