← High Availability: design, failures, and recovery
06 / 6 · 40 MIN

Maintenance and recovery evidence

Decide changes from actual state and close with service and redundancy evidence.

Concept and mechanism

A maintenance plan depends on prerequisites that can change before the window starts. In Kubernetes, a PDB can limit voluntary evictions that respect it, but does not prevent physical node failure or direct Pod deletion. An existing involuntary failure reduces headroom for new evictions. With three replicas, minAvailable=2, and one unavailable replica, removing another healthy replica would leave only one. Bypassing the block through deletion does not restore missing availability. Reassess state, capacity, and communication. Authorization for a window based on three healthy replicas does not establish acceptance of an already-degraded scenario. Restoring the replica or deferring the change may be necessary.

Guided application

Also plan termination: preStop consumes the Pod grace period rather than receiving an additional complete window. Draining, in-flight operations, and shutdown must fit validated behavior. Force-delete removes the API object without confirming that the process stopped, so it is not isolation evidence. For a resilience exercise, define a hypothesis, functional metrics, scope, stop conditions, and recovery. Control impact and record timings and outcomes. In a fictional handover example, green nodes provide only part of the evidence: include consumer operations, integrity, restored redundancy, and owned actions. These examples are decision exercises; no failures, promotions, or cluster changes were executed. Acceptance should preserve the difference between a successful component action and a complete service outcome.

IN PRACTICE

PDB minAvailable=2, two available replicas: there is no headroom to remove another healthy replica.

Common pitfalls

PDB as universal protection; bypass as recovery; hook with separate clock; green dashboard as acceptance.

Related topics: Objectives and service impact · Failure domains and residual capacity · Quorum and writer isolation

Take this idea with you

Validate prerequisites, control the change, and demonstrate the complete outcome.

Create account

Reference: Controlled resilience experiments and evidence · BigSavant HA 2026-09; Pacemaker 3.0, etcd 3.6, PostgreSQL 18 and selected Kubernetes/AWS behavior