1. Choose the failure the design must withstand
The fictional requirement says positions queries must continue after losing a zone. Before drawing replicas, turn that statement into an observable exercise: which operations must continue, with what data, latency and capacity? Losing a process, host, zone or region affects different scopes. A VM placed in a zone remains an instance in that zone; selecting a location does not automatically create another available instance. Draw the components surviving the chosen event and mark those that stop responding. Then follow a complete request from client to data. If it requires a lost component, redundancy elsewhere does not yet demonstrate the intended continuity. Make the failure boundary explicit before accepting a diagram as resilience evidence.
2. Relate placement to correlated failures
Availability sets distribute VMs across fault and update domains to reduce certain shared risks and coordinate maintenance. They should not be presented as equivalent to a design distributed across zones for a zonal outage. In the fictional scenario, two VMs were logically separated but the team never confirmed the isolation type. The architect records the actual configuration and the event it covers. They also compare proximity and latency needs with separation between failure domains. Placing every component close together may support one requirement while concentrating another risk. The decision should show that trade-off, accepted limitations and the exercise that will verify behavior. Counting two VMs alone does not characterize their independence or their application’s ability to survive the chosen event. Across subscriptions, confirm physical mapping before comparing logical zone numbers: zone 1 in one subscription can correspond to a different physical location in another.
3. Size survivors using real demand
Three zones do not guarantee sufficient capacity after losing one. In the exercise, measured usable capacities are 120, 100 and 100 operations per second; agreed peak demand is 180 and required reserve is 20. Losing the largest zone leaves 200, exactly the model’s required total. If reserve rises to 21, the same design fails. These values are fictional and assume additive capacity with distributable work. A real system needs measurements of queues, state, dependencies and load-balancing limits. Do not count capacity not yet created as immediately available without proving provisioning, quotas and startup time. Exercising only during low demand can hide the shortfall that appears precisely during close. Record what each capacity figure includes and which bottlenecks it excludes.
4. Confirm SQL availability configuration
A managed service includes availability mechanisms, but the architect must still select a configuration covering the requirement. In Azure SQL Database, zone redundancy depends on the applicable option, tier and support. A backup in another region does not establish immediate database access after a zone fails. In the fictional case, the diagram says only managed SQL and the team assumes the matter is resolved. Acceptance requires actual configuration, an exercise of client behavior and comparison against the service objective. Connections may need to be re-established and in-flight work may require appropriate handling. Database availability and correct completion of a business operation need related but distinct evidence. Avoid treating a service label as proof of the application behavior promised to users. For example, Basic in the DTU model does not provide zone redundancy. For a read that failed with a transient connection error, establish a fresh connection before retrying under a bounded policy; writes also require transaction-outcome handling.
5. Relate Cosmos DB regions, consistency and clients
For a Cosmos DB design, document where writes are accepted, which regions the client prefers and the chosen consistency. Multiple regions alone do not answer what happens to writes during a failure. A planned region-change procedure assumes different conditions from recovery when a region is unavailable. In the fictional scenario, a previous exercise changed the write region while both were healthy; the team presents it as evidence of outage recovery. Request an exercise matching the actual event, including SDK configuration and the state of observed data. The business must understand the trade-off between continuity and potentially unreplicated changes. Avoid promising the same outcome for every combination of consistency, topology and failover operation. Evidence must match the exact configuration proposed for production.
6. Distinguish local redundancy and geographic storage copies
ZRS distributes copies across zones in the primary region. GRS and GZRS add asynchronous geographic replication, with different protection inside the primary region. Secondary read access must also be selected when the design requires it; do not infer simultaneous writing in both regions. In the fictional case, a consumer tries writing to the secondary endpoint during an outage and the team discovers that the runbook confused reading with service promotion. Describe the recovery operation and required client changes. Also consider recently accepted data that may not have reached the geographic copy. Infrastructure redundancy does not replace protection against an incorrect logical change propagated to the copies. State which failure each mechanism addresses and what recovery evidence is still required.
7. Follow dependencies through to the business result
The front end survives in two zones but depends on an internal configuration service located only in the lost zone. Requests still fail. This fictional case shows why the inventory must include identity, name resolution, certificates, secrets, networking and data alongside the VMs visible in the main diagram. For each dependency, record the effect of unavailability and whether the service can continue using already available information. A retry also needs semantics: when a write outcome is unknown, blind repetition can duplicate an instruction. The application owner must define operation identification, outcome lookup and replay behavior. Measuring healthy hosts alone leaves the consumer experience unproven. Follow the request and its state transitions until the promised business result is observable.
8. Prepare a failure-based acceptance matrix
The final worksheet should connect failure, detection, routing, capacity, data state, client outcome and return to normal operation. In the fictional exercise, each team fills in only its component; the PM identifies a gap between database recovery and batch resumption. The integrated exercise adds that transition and observes the functional outcome instead of declaring a step complete merely because no alarm appeared. Also record who can initiate recovery outside working hours and what evidence authorizes returning to normal configuration. A successful planned change and an isolated restore remain useful, but should support only conclusions they actually demonstrate. Final acceptance should name outstanding limitations and the person responsible for resolving them. This keeps the architecture decision connected to operational ownership.
# Fictional additive capacity model: no Azure calls or availability prediction.
def survives_one_zone(capacity, peak, reserve):
if not capacity or peak < 0 or reserve < 0:
raise ValueError("nonempty zones and nonnegative demand required")
if any(value < 0 for value in capacity.values):
raise ValueError("negative capacity")
total = sum(capacity.values)
return all(total - lost >= peak + reserve for lost in capacity.values)
assert survives_one_zone({"a": 120, "b": 100, "c": 100}, 180, 20)
assert not survives_one_zone({"a": 120, "b": 100, "c": 100}, 180, 21)
assert not survives_one_zone({"a": 300}, 180, 20)
assert survives_one_zone({"a": 210, "b": 210}, 180, 20)
assert not survives_one_zone({"a": 160, "b": 80, "c": 80}, 180, 0)
assert survives_one_zone({"a": 160, "b": 80, "c": 80}, 150, 10)
print("six fictional surviving-capacity checks passed")
Fictional case: three zones support normal load, but losing the largest leaves insufficient capacity during close. The team adjusts capacity and exercises the complete client path before handover.
Common pitfalls
Counting replicas without checking isolation; ignoring surviving capacity; confusing planned change with outage recovery; inferring application continuity from component availability.
Related topics: Failure domains · Capacity and degradation · Recovery and operational acceptance
High availability must survive the chosen failure with sufficient capacity, data and dependencies to deliver the promised user outcome.
Reference: Reliability in Azure Virtual Machines · AZ-305 objectives 2026-04-17