Concept and mechanism
Counting replicas is insufficient for evaluating continuity. Identify what can fail together: host, rack, network, zone, control mechanisms, and application dependencies. Two instances in different zones remain vulnerable if both need one database without an alternative. Distribution should cover the functional path and data access. Then calculate residual capacity. If each of two zones supports 600 requests per second and the agreed peak is 900, losing one leaves only 600. The normal 1,200 does not meet the requirement to maintain peak load during that failure. These figures are original examples rather than vendor specifications.
Guided application
Static stability reduces reliance on creating resources during a crisis: surviving capacity should already support the defined behavior. This has a cost to compare with impact and accepted degraded modes. Autoscaling can help, but needs time, capacity, and available control mechanisms. In Kubernetes, topology spread constraints can distribute Pods; with DoNotSchedule, an incompatible placement can remain Pending. Relaxing to ScheduleAnyway changes the guarantee rather than creating resources. In a fictional etcd cluster with two voters in one zone and one in another, losing the two-voter zone loses the majority. The manager should request evidence for the specific failure and surviving load rather than accept a diagram containing multiple boxes.
Normal capacity of 1,200 requests/s can become 600 when one zone fails.
Common pitfalls
Replicas as independence; normal total as post-failure headroom; instantaneous autoscaling; preference as guarantee.
Related topics: Objectives and service impact · Quorum and writer isolation · Replication, promotion, and redundancy
Demonstrate that the complete path and capacity survive the defined failure.
Reference: Static stability and capacity during failure · DR HA 2026-09; Pacemaker 3.0, etcd 3.6, PostgreSQL 18 and selected Kubernetes/AWS behavior