← Kubernetes: operate workloads and recover services
06 / 6 · 40 MIN

Diagnosis and maintenance

Collect failure evidence and plan maintenance without bypassing availability.

Concept and mechanism

Start an incident with affected context, namespace, workload, and time interval. The same name can identify different objects; replacing a Pod changes the instance and may lose local evidence. Inspect conditions, events, restarts, and logs for the correct container. CrashLoopBackOff describes waiting between restarts rather than the original cause. Where available, logs --previous retrieves the previous terminated instance of that container rather than the service’s entire history. Compare error, configuration, and recent change before repeating a destructive action. Preserve timestamps and keep hypotheses separate from facts. A Running Pod may still lack readiness or fail to produce the result the business expects.

Guided application

During planned maintenance, a PodDisruptionBudget limits applicable voluntary disruption through the Eviction API. It does not prevent involuntary node failures or protect against every direct deletion. A Deployment rollout uses its own availability settings. In an original exercise, four replicas exist, minAvailable is three, and one is already unavailable: no allowance remains for another voluntary eviction in that state. Investigate the unavailable replica or create validated capacity before continuing the drain. Removing protection merely to complete maintenance may transfer risk to the service. At RUN handover, deliver health criteria, limits, dependencies, recovery, and ownership; validate a representative operation and backlog rather than only command completion.

IN PRACTICE

Four replicas, three healthy, minAvailable=3: additional maintenance requires restoring headroom.

Common pitfalls

CrashLoopBackOff as cause; previous as a complete archive; PDB as universal protection.

Related topics: Workloads and desired state · Traffic and probes · Resources and scheduling

Take this idea with you

Recover function and respect observable availability limits.

Create account

Reference: Disruption budgets and eviction · Kubernetes v1.37 concepts; current official documentation consulted 2026-09-30; cluster versions and plugin capabilities must be confirmed