← Back to catalogue
Professional assessment

High Availability: design, failures, and recovery

Six lessons, 30 questions, and six cases on availability objectives, capacity, quorum, fencing, replication, probes, and maintenance.

BigSavantAvailable
BigSavantHA6 lessons
Internal assessment of 24 decisions in 60 minutes. Availability objectives, residual capacity, quorum, fencing, replication, probes, retries, and maintenance. No external exam or certification.

Objectives and progression

A six-module technical course with fictional APS and infrastructure management examples. Learn to measure impact, identify shared dependencies, calculate residual capacity and quorum, assess fencing and replication, and prepare maintenance with demonstrable recovery. Includes selected Pacemaker 3.0, etcd 3.6, PostgreSQL 18, and Kubernetes behavior, primary references, and an internal assessment of 24 decisions in 60 minutes.

Audience: APS L2/L3, infrastructure, SRE, systems administration teams, and technical managers.

Prerequisites: Networking, storage, and service fundamentals; examples state relevant conditions and versions.

300 estimated study minutes

  • Define availability through service outcome, with explicit population, window, and criteria.
  • Assess shared dependencies, placement, and capacity after the expected failure.
  • Distinguish majority decisions, suspected failure, and proven isolation.
  • Relate commit acknowledgement, available data, and safe recovery of the primary role.
  • Choose signals and attempts that support recovery without amplifying failure.
  • Decide changes from actual state and close with service and redundancy evidence.

Modules

  1. Objectives and service impact
  2. Failure domains and residual capacity
  3. Quorum and writer isolation
  4. Replication, promotion, and redundancy
  5. Health, traffic, and retry pressure
  6. Maintenance and recovery evidence

Continue learning

Technical assessments

References and version

DR HA 2026-09; Pacemaker 3.0, etcd 3.6, PostgreSQL 18 and selected Kubernetes/AWS behavior

What you will explore

0 / 6

Learning is also trying.

Original explained questions, flashcards, and scenarios to apply the concepts.

Practice
This module covers foundations. It is not a complete certification course or a full simulation of the official exam.

Assessment topics

Objectives and service impact—
Failure domains and capacity—
Quorum and isolation—
Replication and promotion—
Health, traffic, and retries—
Maintenance and recovery evidence—