← High Availability: design, failures, and recovery
05 / 6 · 40 MIN

Health, traffic, and retry pressure

Choose signals and attempts that support recovery without amplifying failure.

Concept and mechanism

Kubernetes probes serve different purposes. Readiness indicates whether a container is ready for traffic from matching Services; failure is not itself a restart instruction. Liveness can cause a restart once failure exceeds configured tolerance. A startup probe accommodates lengthy initialization and gates liveness and readiness until success. Confirm timings and conditions appropriate to the process. An application that needs two minutes to load data should not be repeatedly restarted before finishing that work. Making liveness very slow throughout the container lifetime can hide later deadlocks; separating startup preserves useful detection during operation. Probe design should follow the recovery action that actually helps the application.

Guided application

In a fictional example, every replica depends on a slow backend. A liveness check that fails because of that dependency restarts healthy replicas, reduces capacity, and increases load on survivors. Indiscriminate retrying worsens the cycle. Define eligible errors, limits, and backoff with jitter, and inspect retries already performed by other layers. If three layers make three total attempts each and every attempt fails, the model can produce 27 calls to the final dependency. For mutations, timeout does not establish absence of effect: validate idempotency or reconcile the outcome. Jitter spreads attempts over time but does not eliminate business duplicates. Recovery should be measured through sustained service outcomes rather than one passing probe.

IN PRACTICE

Three layers with three total attempts can amplify one request to 3 × 3 × 3 = 27.

Common pitfalls

Readiness as restart; liveness depending on any delay; unlimited retries; jitter as idempotency.

Related topics: Objectives and service impact · Failure domains and residual capacity · Quorum and writer isolation

Take this idea with you

Separate readiness and recovery, and bound additional load during failures.

Create account

Reference: Kubernetes startup, readiness and liveness probes · DR HA 2026-09; Pacemaker 3.0, etcd 3.6, PostgreSQL 18 and selected Kubernetes/AWS behavior