High Availability: design, failures, and recovery
Six lessons, 30 questions, and six cases on availability objectives, capacity, quorum, fencing, replication, probes, and maintenance.
Objectives and progression
A six-module technical course with fictional APS and infrastructure management examples. Learn to measure impact, identify shared dependencies, calculate residual capacity and quorum, assess fencing and replication, and prepare maintenance with demonstrable recovery. Includes selected Pacemaker 3.0, etcd 3.6, PostgreSQL 18, and Kubernetes behavior, primary references, and an internal assessment of 24 decisions in 60 minutes.
Audience: APS L2/L3, infrastructure, SRE, systems administration teams, and technical managers.
Prerequisites: Networking, storage, and service fundamentals; examples state relevant conditions and versions.
300 estimated study minutes
- Define availability through service outcome, with explicit population, window, and criteria.
- Assess shared dependencies, placement, and capacity after the expected failure.
- Distinguish majority decisions, suspected failure, and proven isolation.
- Relate commit acknowledgement, available data, and safe recovery of the primary role.
- Choose signals and attempts that support recovery without amplifying failure.
- Decide changes from actual state and close with service and redundancy evidence.
Modules
- Objectives and service impact
- Failure domains and residual capacity
- Quorum and writer isolation
- Replication, promotion, and redundancy
- Health, traffic, and retry pressure
- Maintenance and recovery evidence
Continue learning
References and version
DR HA 2026-09; Pacemaker 3.0, etcd 3.6, PostgreSQL 18 and selected Kubernetes/AWS behavior
- Service level objectives and user-facing indicators · 2026-09-30
- Availability and observation-window arithmetic · 2026-09-30
- Deploy the workload to multiple locations · 2026-09-30
- Static stability and capacity during failure · 2026-09-30
- Controlled resilience experiments and evidence · 2026-09-30
- Control and limit retry calls · 2026-09-30
- Pacemaker 3.0 fencing and isolation · 2026-09-30
- etcd 3.6 quorum and fault tolerance · 2026-09-30
- PostgreSQL 18 standby replication and commit acknowledgement · 2026-09-30
- PostgreSQL 18 failover and restoration of redundancy · 2026-09-30
- Kubernetes startup, readiness and liveness probes · 2026-09-30
- Kubernetes disruptions and PDB scope · 2026-09-30
- Kubernetes topology spread constraints · 2026-09-30
- Kubernetes graceful and forced Pod termination · 2026-09-30
What you will explore
0 / 6Objectives and service impact
Define availability through service outcome, with explicit population, window, and criteria.
Failure domains and residual capacity
Assess shared dependencies, placement, and capacity after the expected failure.
Quorum and writer isolation
Distinguish majority decisions, suspected failure, and proven isolation.
Replication, promotion, and redundancy
Relate commit acknowledgement, available data, and safe recovery of the primary role.
Health, traffic, and retry pressure
Choose signals and attempts that support recovery without amplifying failure.
Maintenance and recovery evidence
Decide changes from actual state and close with service and redundancy evidence.