← AZ-305: Azure architecture and production decisions
14 / 23 · 100 MIN

Blob Storage: recovery, retention and failover

Choose protection from the type of loss and prepare reconciliation, recovery and restored redundancy.

1. Classify the loss before starting recovery

In a fictional fund-processing service, a file can disappear through deletion, be replaced with an incorrect version or become inaccessible during a regional outage. These incidents call for different decisions. A geographic replica helps with regional unavailability but can receive the same unwanted change as the source. Before selecting an operation, record the affected object, the last business-accepted version, the suspected change interval and service availability. Separate restoring access from restoring correct data. An endpoint responding again can still serve an incorrect file. During the project, require acceptance criteria for both outcomes and assign recovered-content validation to the business, supported by technical integrity checks and traceability. The incident record should explain what evidence supports the selected recovery path.

2. Identify the deleted unit

Blob soft delete and container soft delete protect different units. In an exercise, an operator deletes an entire container although the team rehearsed only individual blob recovery. The runbook must recognize that difference before running commands. If previous blob versions exist, identify which represents the intended state and how it will become current. With versioning enabled, undelete alone does not promote an earlier version to the current version; recovery can require copying it for that purpose. Record its version identifier and validate size, hash and business content. Avoid simply selecting the newest version when the incident is an incorrect overwrite. Rehearse recovery permissions without granting unnecessary permanent support access, and retain enough evidence to explain the selected version during the incident review.

3. Prepare the recovery point before the incident

Point-in-time restore has prerequisites and compatibility limits: enabling it after discovering corruption is insufficient. Configuration requires soft delete, change feed and versioning; protection does not create history predating enablement. Check the supported account type, namespace and tiers in current documentation. The operation can block access to the ranges being restored, so include that interruption in the plan. In a fictional load that changed thousands of positions, the manager asks the team to bound the affected set and choose a UTC instant preceding the load. Approval includes the impact on legitimate later changes, evidence that the recovery point exists and post-recovery validation. Do not reuse this procedure to recover container deletion. The rehearsal should produce observable outcomes rather than only a successful configuration screenshot.

4. Interpret Last Sync Time as a certainty boundary

During a regional outage, asynchronous replication creates an uncertainty window. Last Sync Time distinguishes earlier writes, whose presence at the secondary is guaranteed by the service, from later writes that might not yet be there. It does not mean every later write was lost. In a fictional example, the last confirmed instant is 14:00 UTC and operations are recorded until 14:07. The team marks those operations for reconciliation, preserves identifiers and checks recovered outcomes before repeating work. The Python model uses integer minutes and a fictional operation ledger to separate sets; it neither queries Azure nor measures actual loss. The failover decision needs an owner who accepts the known risk and defines treatment of unresolved outcomes. Repeating every uncertain instruction could create a second business error.

5. Finish failover by restoring protection

After customer-managed unplanned failover, the new primary has local redundancy. Restoring service does not finish the resilience plan: restoring geographic redundancy, monitoring replication and accounting for cost remain necessary. A planned exercise has different conditions, including region availability and compatibility with enabled features. Do not disable data protection merely to make a rehearsal succeed without assessing the exposure interval created. At handover, keep a separate task for confirming the final protection state, with an owner and evidence. If the service temporarily operates with reduced resilience, give that exception an expiry and follow-up. Returning to the original region is another planned operation with its own integrity, access and decision criteria. Availability reports should show outstanding protection work alongside the service-restoration milestone.

6. Align retention, cost and recovery evidence

Operational retention and business-required retention must align with deletion policies. A FinOps team might propose deleting older versions to reduce cost while a control team needs files from an earlier closing cycle. The project should reconcile both requirements using concrete dates and recovery examples. Do not assume the point-in-time restore window automatically removes all older versions: review lifecycle rules applied to versions and the effect of soft delete. In a rehearsal, select a known object, change it, recover the intended state and compare business evidence. Document measured time and rehearsal limitations. A successful small execution does not demonstrate recovery time for the entire production volume. The acceptance record should state the tested scale, remaining uncertainty and the owner of any follow-up capacity exercise.

writes = [("A", 837), ("B", 839), ("C", 843), ("D", 847)]
last_sync = 840 # fictional UTC minute of day
confirmed = {key for key, minute in writes if minute < last_sync}
uncertain = {key for key, minute in writes if minute >= last_sync}
assert confirmed == {"A", "B"}
assert uncertain == {"C", "D"}
assert confirmed.isdisjoint(uncertain)
assert confirmed | uncertain == {key for key, _ in writes}
recovered = {"A", "B", "C"} # hypothetical observed result, not inferred from LST
assert uncertain - recovered == {"D"}
print("five fictional reconciliation checks passed; no Azure loss measured")
IN PRACTICE

Fictional case: the region fails at 14:07 and Last Sync Time is 14:00. Service is restored, but seven minutes of instructions require reconciliation. Incident closure needs confirmed outcomes and a plan to restore geographic redundancy.

Common pitfalls

Confusing replication with history; using blob protection to eliminate container-deletion risk; treating all writes after Last Sync Time as certainly lost; closing without restoring redundancy.

Related topics: RPO and RTO · Reconciliation and idempotency · FinOps and retention

Take this idea with you

Recovery requires restoring correct data, confirming business outcomes and restoring agreed protection.

Create account

Reference: Blob point-in-time restore · AZ-305 objectives 2026-04-17

Azure is a trademark of the Microsoft group of companies. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Microsoft. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.