1. Recover an application with its dependencies
A recovery plan organizes machines and actions into a recovery sequence. Start from application dependencies: identities, DNS, database, middleware, files and traffic entry. Machines within one group can start in parallel; separate groups impose order. However, a running VM does not prove its service meets the expected contract. Add readiness checks and appropriate actions before opening consumer access. In a fictional case, the database starts before middleware but required logins have not been validated. Startup order is correct and the service remains unavailable. Assign each check to a team and define what blocks progression to the next stage. This turns orchestration into an application recovery process rather than a list of successfully started virtual machines.
2. Rehearse without creating a second active producer
Test failover exercises a copy while the source continues operating. Prepare the test network so the copy cannot send real instructions, emails or files to production consumers. Network, DNS and dependency isolation belong in rehearsal design. In the fictional example, a replicated VM contains an automatically starting scheduler and file-integration credentials. Merely starting it can execute unintended work if the real destination remains reachable. Define simulated or controlled destinations, fictional data and functional tests compatible with isolation. Also record which real paths the rehearsal excludes and which need separate validation. Lack of production impact should follow verified controls rather than a hope that nobody uses the copy. Include automatically triggered work when checking the test boundary.
3. Choose the point from acceptable loss and recovery time
Recovery-point selection involves trade-offs. An already-processed point can avoid additional work before recovery starts; requesting the latest point can require processing data already received by the service. An application-consistent point can be older than another available point. Do not confuse an option’s name with a universal RPO or RTO guarantee. Compare data time, required consistency and estimated recovery and validation duration. The local exercise uses fictional timings supplied by the team. If no candidate meets every limit, the function returns an empty list. That is a useful conclusion: the design needs improvement or a formal trade-off decision rather than presenting the least-bad option as automatic compliance. Actual recovery evidence must come from the workload and the selected service configuration.
4. Distinguish retention, soft delete and immutability
Retention defines how long recovery points should be kept; soft delete provides an opportunity to recover data after deletion; immutability blocks operations that could remove protected points. These capabilities are not synonyms. A vault with immutability enabled can still have a reversible setting; locked state makes that setting irreversible. Before locking, validate policy, cost and permitted-operation effects. Do not generalize vault protection to every operational backup of blobs, files or disks either. Confirm feature scope for the backup type in use. In the fictional project, the requirement is to withstand malicious deletion and retain usable recovery. Keeping data that nobody can restore does not meet the second part. Acceptance therefore needs both protection evidence and a successful authorized restoration path.
5. Preserve separation of duties during recovery
Resource Guard adds authorization for critical Azure Backup operations. Separation loses value if the vault administrator retains permanent privileges allowing them to approve or execute every protected operation alone. Define a distinct owner, required access and emergency handling. Review protected and excluded operations for the vault type instead of assuming identical defaults everywhere. The fictional scenario requires a retention reduction to pass the agreed control while urgent restoration still has a known path. Test authorization with separated roles and document how approval is obtained outside working hours. The control should reduce the risk from one compromised identity without making recovery depend on an unavailable individual. The handover needs practical contacts and permissions, not merely a diagram showing two teams.
6. Close the rehearsal with outcomes and cleanup
A rehearsal ends with functional validation, recorded gaps and cleanup of created copies. Changes made inside a test-failover VM must not be treated as fixes automatically returned to the source. If testing reveals missing configuration, correct the appropriate controlled source and repeat the rehearsal. Preserve evidence before cleanup: selected point, timings, outcomes, isolation limits and open tasks. Backup-protection documentation can change by region and vault type; where official descriptions disagree, record uncertainty and validate effective state before approving the design. Do not use that ambiguity to declare universal capability. The final decision should state what was demonstrated and what remains before complete-service recovery can work in the real context. Rehearsal cleanup should not erase the evidence needed to close those gaps.
points = [
{"id": "processed", "age": 8, "restore": 3, "app_consistent": False},
{"id": "latest", "age": 2, "restore": 8, "app_consistent": False},
{"id": "app", "age": 12, "restore": 4, "app_consistent": True},
]
def eligible(max_age, max_total, validation, require_app=False):
return [p["id"] for p in points
if p["age"] <= max_age and p["restore"] + validation <= max_total
and (not require_app or p["app_consistent"])]
assert eligible(5, 6, 2) == []
assert eligible(5, 10, 2) == ["latest"]
assert eligible(10, 6, 2) == ["processed"]
assert eligible(15, 6, 2, True) == ["app"]
assert eligible(5, 10, 2, True) == []
assert eligible(15, 6, 3, True) == []
print("six supplied recovery-point checks passed; no Azure recovery performed")
Fictional case: the latest point meets RPO, but processing it and validating the application exceeds RTO. The processed point starts faster but is too old. The project keeps the gap open and improves the design.
Common pitfalls
Treating a started VM as a ready application; allowing test jobs to contact production; equating the latest point with every recovery objective; treating enabled immutability as locked.
Related topics: Recovery dependencies · RPO and RTO · Separation of duties
Recovery needs a suitable point, authorized execution and service evidence with coherent protection and isolation.
Reference: Test failover to Azure · AZ-305 objectives 2026-04-17