← High Availability: design, failures, and recovery
08 / 8 · 60 MIN

Election, rejoining, and maintenance

Interpret cluster failures and recovery without confusing available processes, votes, and consumer recovery.

Stop a follower and observe the remaining service

After identifying the leader, the lab stops a follower belonging to the cluster. The other two members acknowledge a write containing after-one-stop. Stopping does not remove the member from membership: three voters remain configured and two are available. This distinction matters during maintenance. An inventory of live processes does not replace membership and communication between voters. The exercise demonstrates majority progress in this local case; it does not promise no impact for every client. A consumer knowing only the stopped endpoint may remain unavailable while a direct operation against a healthy member works. Diagnosis should retain both observations.

Rejoin the same member with its data

The follower restarts with its original directory without creating another cluster or changing membership. The lab waits until it can read on that member the value committed during its absence. This is more useful evidence than merely seeing the process start. The procedure applies to a temporarily stopped member that was not removed; a permanently removed member has different rules and must not be reintegrated by analogy. In a runbook, distinguish restart, replacement, and disaster recovery. Before closing maintenance, confirm identity, state, operations, and remaining headroom for the next failure. Starting with the correct name does not prove all of this.

Leader election and client recovery

In another group, the lab abruptly terminates only the process identified as leader. The two surviving members converge on a different leader and acknowledge after-leader-stop. The former leader returns and catches up with that state. The exercise does not measure a contractual recovery bound. It has a deadline for failure if the condition is not observed, but that deadline is not an SLA. A real application may need to reconnect, renew sessions, discover endpoints, and handle unanswered operations. Completed election and recovered batch are therefore distinct milestones. In a committee, retain timing and scope for each observation and avoid extrapolating from a laptop to a distributed architecture under load.

Restore quorum and reconcile uncertain operations

When two voters are stopped, the exercise sends a put to the remaining process and receives no acknowledgement within the deadline. It then restarts one stopped member, observes a functioning majority, and reads the key from that attempt. In the recorded run, the result was absence. The code allows presence or absence because client timeout alone does not determine final outcome. The lesson is reconciliation, not memorizing this execution output. The exercise then acknowledges a new write and all three members return, without forced membership changes. For a service with external effects, define how to resolve operation identity and outcome before a retry that might duplicate work.

Learners and maintenance prerequisites

Learner examples are documentary analysis, without membership changes executed in this lab. An unpromoted learner receives state but does not yet count as a voter. Thus, three voters and one learner still require two votes. Promotion requires server-checked conditions, including appropriate catch-up. Do not remove another voter merely because the learner process appears on a dashboard. Documentation recommends controlled changes verified one at a time. Quorum loss is also not solved by ordinary member add sent to the sole survivor, because the change needs cluster agreement. Distinguish recovering members, replacing a member, and following a recovery procedure when a majority cannot return.

Acceptance through capacity, dependencies, and function

In a planning example, three zones offer 400 requests/s each and load is 750. Losing one leaves 800 and nominal headroom of 50, assuming traffic and dependencies allow that capacity to be used. The calculation does not establish working identity or configuration if those services reside in the lost zone. Acceptance should connect the defined failure to consumer function, capacity, errors, and reconciliation. The local lab provides observations about processes and consensus; it did not execute zone loss, a real network partition, physical fencing, or PostgreSQL promotion. Recorded evidence includes version and script hash, with environment exercises and specialist review still needed for broader conclusions.

configured voters = 3
1 stopped -> 2 live voters, majority possible
2 stopped -> 1 live voter, majority unavailable
learner before promotion -> no additional vote
# Process count alone is not quorum or service acceptance.
IN PRACTICE

Stopping one follower permits progress with two voters; stopping two leaves a reachable process without quorum. Returning a member restores operations, but the timed-out put outcome is reconciled separately.

Common pitfalls

Counting learners as votes; recreating a cluster during restart; treating timeout as absence; measuring only election; using three same-host processes as proof of zone isolation.

Related topics: Quorum, reads, and authority · Residual capacity and failure domains · Fencing and service recovery

Take this idea with you

Safe operation requires knowing which members vote, which state returned, and whether the consumer can complete the agreed function.

Create account

Reference: etcd failure modes · BigSavant HA 2026-09; Pacemaker 3.0, etcd 3.6, PostgreSQL 18 and selected Kubernetes/AWS behavior