← High Availability: design, failures, and recovery
10 / 12 · 60 MIN

Replacement, identity, and acceptance

Reconcile membership and identity after replacement and close the change with consumer evidence and handover.

An exercise with explicit assumptions

This workshop continues the healthy lab cluster: three original voters and a candidate already promoted. The sequence was chosen to observe four votes, one shutdown, and subsequent removal of the old identity. It is not a universal replacement procedure. Reconfiguration documentation also presents replacement through removal followed by addition; the choice depends on state, version, and the applicable procedure. Under majority loss, do not assume the healthy sequence will accept new changes. Distinguish member recovery, planned replacement, and disaster recovery before preparing commands or making deadline commitments to the business.

Stopping and removing are different transitions

The lab selects an original member that is not the leader and stops only its process. Four votes remain configured and three processes remain active, sufficient for the majority of three. A synthetic write confirms progress in this state. Removal of the old identity is then submitted to the cluster. Inventory becomes three voters and the required majority returns to two. Evidence confirms the old identity is absent, the replacement is present, and the synthetic value exists on active members. A timeline of these observations is more useful than simply reporting server retired because it reconstructs headroom at each stage.

Return of a removed identity

The exercise starts the old process again using its own disposable directory. The process recognizes that its identity was removed and exits; membership retains the replacement. This demonstrates why rollback cannot simply mean running the previous binary again. Removal is cluster state. Real recovery needs a procedure compatible with that state and protection for data that investigation may require. The lab accepts no external directories and removes only the temporary data it created. Do not use this exercise as authorization to erase or repurpose directories belonging to an existing installation.

Reconcile before repeating

After replacement, the lab attempts to promote the already voting member again. The command is rejected even though the desired state already exists. This observation separates objective, state, and exit code. An executor that only accepts command success may try to undo correct work to satisfy its script. With an earlier timeout, the problem differs: state is not yet known. In both cases, inspect membership and preserve target identity before choosing the next action. Record who observed the state, when, and what remains uncertain so that a shift change does not restart the sequence based on an assumption.

Accept the consumer and remaining headroom

In the final three-voter cluster, the lab stops a follower, confirms another write, and rejoins that process with its own data. This demonstrates the local observation of one shutdown, not zone loss or recovery of a banking job. A consumer configured only for the retired endpoint may still fail when the cluster already works. Define functional operations, source network, acceptable latency, observation window, and acceptance owner. Record cluster recovery, consumer recovery, and restored redundancy separately. The laptop exercise does not qualify production Linux, security, capacity, or independent physical failure domains.

Handover and closure decision

Prepare a forty-five-minute exercise in English. For fifteen minutes, the team builds a fictional replacement timeline and calculates the majority at every stage. At minute fifteen, it receives an inject: the cluster works, but the reconciliation job targets the old endpoint. Use another fifteen minutes to assign diagnosis, communicate cut-off impact, and define the next update. During the final fifteen minutes, present a handover containing membership, removed identity, evidence, pending consumer, owner, and a boundary for further changes. Close only demonstrated criteria; administrative approval does not turn a still-failing operation into functional recovery.

3 voters + 1 learner -> majority 2
4 voters after promotion -> majority 3
1 voter process stopped -> still 4 configured, 3 live
old identity removed -> 3 configured, majority 2
# Observed local sequence; assess prerequisites for any real change.
IN PRACTICE

Stopping the old member leaves four configured votes and three active processes. Removing its identity leaves three votes. Starting its old data does not undo that removal.

Common pitfalls

Calling a simple restart a rollback; repeating operations after timeout; measuring only cluster health; accepting same-host tests as proof of zone-loss resilience.

Related topics: Failure domains and residual capacity · Quorum and writer isolation · Election, rejoining, and maintenance

Take this idea with you

Close the change with reconciled identity, membership, and functional outcome. Keep limitations and outstanding acceptance criteria visible.

Create account

Reference: etcd runtime reconfiguration · BigSavant HA 2026-09; Pacemaker 3.0, etcd 3.6, PostgreSQL 18 and selected Kubernetes/AWS behavior