← AZ-305: Azure architecture and production decisions
19 / 23 · 100 MIN

App Service: deployment, networking and controlled recovery

Prepare slot-based changes, validate identity and connectivity and distinguish code rollback from data recovery.

1. Prepare the configuration that will actually run

In a fictional operations portal, the new release works in staging but fails after promotion. The team tested code and sample data without checking identity and dependencies under production configuration. In App Service, not everything follows content through a swap: managed identities and VNet integration remain associated with the slot. Some application settings and connection strings can be marked as slot-specific. Maintain a short matrix with component, expected behavior, owner and access evidence. Rehearse the final combination as well as preparation. A successful connection using an administrative development account does not demonstrate that the production identity can read the secret or execute the required operation. Use the intended identity and an authorized, controlled check when collecting acceptance evidence.

2. Warm up and evaluate readiness without business effects

The warm-up endpoint should prepare the application without generating financial instructions, real emails or duplicate jobs. In the fictional example, the home page responds even when the database connection is unavailable. Accepting that response lets the new version proceed without proving the critical dependency. Define explicit readiness criteria and confirm warm-up configuration instead of inferring readiness from any HTTP response. Health Check also needs a path with suitable semantics: 2xx responses indicate health; login redirects are not success. Do not let an optional dependency remove every instance when the core service can still operate. Separate minimum health, degraded features and release authorization using signals the team can explain. Test these decisions with controlled failures before relying on them during a deployment window.

3. Design ingress, egress and name resolution

Draw the two traffic directions separately. A private endpoint provides private ingress; VNet integration serves outbound paths. The fictional application must receive requests from an internal network and query a database at a private address. Demonstrating ingress alone does not prove the second path. Record the client hostname, expected resolution, route and authorization for each dependency. If public access is prohibited, explicitly configure and check that restriction. Private deployments also require the CI/CD agent to reach the appropriate endpoint; application-name resolution does not replace SCM resolution. A test from an administrator’s laptop can use different DNS and routes from the agent or application. Collect evidence at the actual execution points instead of treating one successful connection as proof for every caller.

4. Limit exposure and define when to stop

Progressive promotion needs an observation criterion, not just a traffic percentage. For the fictional portal, track technical errors, completed instructions and latency for a representative operation. Define the sample and observation period needed before increasing exposure. A few error-free calls do not demonstrate capacity for the daily closing peak. The team should know who can stop the change and what happens when a metric crosses the agreed limit. Consider the effects of stopping, including accepted work and consumers that received the new contract. Also record telemetry failures: missing data must not automatically authorize the next stage. These criteria belong in the change plan and RUN handover. Use business outcomes alongside infrastructure signals to decide whether the release is functioning as intended.

5. Preserve compatibility during the rollback window

Switching code does not automatically undo changes to an external database. In the fictional case, the new release removes a column that the earlier version still reads. Returning to the old code leaves the service unable to function. Plan compatible evolution: add the new structure, prepare reads and writes, migrate data under control and remove the old structure only after the agreed rollback window. Test both versions against the expected state during that window. If data restoration is required, consider valid instructions received since the recovery point and how to reconcile them. A procedure restoring availability while losing accepted work can fail the business requirement. The local model below shows only selected field movement for an exercise; it does not simulate Azure’s complete swap algorithm.

6. Separate release changes from regional recovery

Slots within an application are not a second recovery region. For a regional-outage requirement, identify another deployment, recoverable data, traffic entry and every required dependency. A diagram with two applications remains incomplete if both depend on one secret, endpoint or manual process unavailable during the incident. In the fictional project, RUN accepts handover only after observing a functional operation at the destination and a return procedure consistent with the data. Record available capacity, measured time, reconciliation decisions and responsibilities. Zone redundancy addresses a different failure scope and should be evaluated separately. Final delivery includes evidence of deployment, recovery and daily operation; a second copy of code does not independently demonstrate those outcomes. Track untested dependencies as open acceptance items with accountable owners.

from copy import deepcopy

def exercise_swap(left, right):
 # Selected steady-state fields only; not an Azure swap implementation.
 a, b = deepcopy(left), deepcopy(right)
 a["code"], b["code"] = b["code"], a["code"]
 return a, b

production = {"code": "v1", "identity": "prod-id", "sticky_db": "prod-db"}
staging = {"code": "v2", "identity": "stage-id", "sticky_db": "stage-db"}
after_prod, after_stage = exercise_swap(production, staging)
assert after_prod["code"] == "v2"
assert after_stage["code"] == "v1"
assert after_prod["identity"] == "prod-id"
assert after_prod["sticky_db"] == "prod-db"
assert production["code"] == "v1" # inputs preserved
assert exercise_swap(after_prod, after_stage) == (production, staging)
print("six selected-field checks passed; no Azure swap or database rollback performed")
IN PRACTICE

Fictional case: staging passes tests, but the production identity cannot access the secret. The owner stops promotion, validates required access and repeats the rehearsal without indiscriminately broadening permissions.

Common pitfalls

Assuming identity follows code; using a warm-up endpoint with business effects; confusing private ingress with egress; promising data reversal through a swap.

Related topics: CI/CD and promotion criteria · Data compatibility · Regional recovery

Take this idea with you

A safe change depends on the code, configuration, identity, network and data combination that will actually run.

Create account

Reference: Set up staging environments in App Service · AZ-305 objectives 2026-04-17

Azure is a trademark of the Microsoft group of companies. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Microsoft. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.