← CKA: Kubernetes administration and troubleshooting
07 / 7 · 24 MIN

Nodes and control plane

Investigate local components when the API or node stops responding.

Concept and mechanism

NotReady on a node requires observing conditions, events, kubelet, runtime, and control-plane communication. A localized failure does not itself justify reinstalling the entire cluster. DiskPressure may result from saturated filesystems and ephemeral storage; identify consumers and recovery mechanisms before deleting runtime directories. Logs and images have lifecycles needing management. Indiscriminate cleanup can destroy data and evidence. Mitigation should restore headroom and address growth origin to prevent repetition during the next business window.

Guided application

kubectl top depends on the Metrics API and a working pipeline, often with metrics-server. A responding main API does not establish that the aggregated metrics API is healthy. Static Pods are managed by the kubelet through local configuration; API mirror Pods are not the manifest control source. If an API server stops starting after an edit, local logs and runtime state may remain accessible. For x509 errors, identify certificate, issuer, expiration, and CA model. kubeadm certs check-expiration helps inventory managed certificates, but renewal and reload should respect context, including external CA.

IN PRACTICE

The API server fails after a local change. With authorized node access, compare the manifest with the known version and inspect kubelet/runtime; do not depend exclusively on kubectl against the unavailable API.

Common pitfalls

Deleting in-use directories; confusing the main API with Metrics API; fixing only the mirror Pod; disabling TLS as an expiration solution.

Related topics: Cluster, context, and maintenance · Access, tools, and extensions

Take this idea with you

Retain local diagnostic paths and treat credentials and storage as operational dependencies.

Create account

Reference: Troubleshoot clusters · CKA Kubernetes v1.35