Concept and mechanism
High availability starts with an outcome that matters to the consumer. A powered-on VM can host an API that fails every request. Choose a service indicator such as successfully completed eligible operations, and define success, the measured population, and the observation window. Do not mix a request-based indicator with time-based availability without explaining the difference. In a time-based example, a month of exactly 30 days contains 43,200 minutes. A 99.95% objective with no exclusions permits 0.05% unavailability: 21.6 minutes. The arithmetic does not decide which events should be excluded; that rule must already be agreed.
Guided application
In a fictional APS meeting, the supplier reports full uptime, but reconciliation missed its cut-off. The technical manager should separate infrastructure health, functional processing, and business impact. A global average can also hide an entirely affected location, so segmentation supports diagnosis without silently changing the contractual indicator. During an exercise, record detection, decision, promotion, and consumer recovery. If promotion takes 20 seconds but operations only return after four minutes, those are different milestones. Functional acceptance uses the agreed outcome and preserves the timeline. Planned maintenance is not automatically excluded: apply the definition and retain evidence of observed impact.
99.95% over 30 days without exclusions: 43,200 × 0.0005 = 21.6 minutes.
Common pitfalls
Powered-on VM as functional service; unsegmented average; retrospective exclusions; promotion as complete recovery.
Related topics: Failure domains and residual capacity · Quorum and writer isolation · Replication, promotion, and redundancy
Agree the outcome, measure impact, and distinguish recovery milestones.
Reference: Service level objectives and user-facing indicators · DR HA 2026-09; Pacemaker 3.0, etcd 3.6, PostgreSQL 18 and selected Kubernetes/AWS behavior