1. Draw the sequence before looking for the fault
In a fictional L3 support scenario, the application has not received the change, but the manager sees several green approvals. Start by drawing the sequence: plan construction, condition evaluation, resource access, job execution and service observation. Ask where work stopped and what evidence supports that conclusion. A job that never started has a different problem from a script that exited with an error. Human approval does not establish that the agent authenticated to the application. Use run and artifact identifiers to avoid mixing attempts. Prepare a short record containing state, last observed action, outstanding condition, owner and window deadline. This structure helps the PM coordinate infrastructure, development and production without asking every team to repeat the same investigation. Before rerunning, identify effects already produced and what repetition could duplicate.
2. Know when a value becomes available
Environment inventory calculates needsMigration during Discover. That result did not exist when the plan was expanded. Define the consumer in the plan and use a runtime condition after the producer, instead of expecting a template if to be reconstructed halfway through execution. Another distinction appears inside a script: its macro was already replaced when the step started. A variable-setting command does not retroactively change that text. The following step can consume the updated value. In the exercise, releaseLabel starts as candidate and later becomes approved. Ask the learner to predict the value observed at each point and explain why. Do not use an arbitrary wait to repair a phase difference. Record each value’s origin, calculation time and first eligible consumer. This small table avoids diagnosis based only on YAML’s visual order.
3. Treat outputs as contracts between producers and consumers
For an output to cross jobs, identify the producer, named step, published variable and consumer dependency. In the ordinary same-stage example, Inspect publishes decide.allowRelease and Release depends on Inspect. Across stages, the consuming job’s context includes stageDependencies, the stage name and producing job. Do not copy the same path into deployment jobs without reviewing their specific syntax. Adding Review between Build and Deploy also needs attention: if Deploy needs Build’s original output, declare that relationship instead of relying only on file position. The case uses packageRef to retain assessed-package identity. Recalculating it from the current main tip can identify different content. In handover, explain which run produced the value and what happens if it is missing. An empty string should not silently become permission to deliver a different package.
4. Distinguish condition, reachability and functional outcome
A condition is evaluated before starting the unit it belongs to. A job cannot base its own start on a result that only its first step would produce. Also, always on a child does not start a parent stage marked Skipped. In the batch scenario, Prepare paused the consumer and restoration was placed inside that unreachable stage. Support needs to confirm external state and perform authorized recovery rather than assume configuration equals completed action. In the local exercise, Tests must finish Succeeded, while Docs may finish Succeeded or Skipped. The contract is deliberately strict and excludes SucceededWithIssues. That choice belongs to the scenario; it is not a universal description of succeeded. If a task uses continueOnError, continued flow also does not turn a failed test into a functional pass. Retain both pieces of information in reporting.
5. Check all access conditions
A stage can consume an environment, service connection and other protected resources. Build a list of resources actually used and pending checks on each. In the practical case, the environment is approved but the service connection still awaits a decision. Repeating the vote on the first resource does not resolve the second. Pipeline authorization is also explicit: copying YAML into a new pipeline does not automatically transfer the old pipeline’s access. Another investigation may find Branch control rejecting a linked artifact run’s origin even though deployment runs on main. Confirm the permitted fully qualified reference and the run actually consumed. Finally, do not confuse configuration changes with updates to an ongoing evaluation: the approver list is fixed when checks begin. Plan accountable people’s availability before the window and confirm how a new attempt will be assessed.
6. Separate receipt, decision and execution window
In an asynchronous integration, the initial HTTP response acknowledges receipt; the callback communicates the access decision. If the external service finishes its work without sending the decision, the pipeline still lacks that evidence. Record the instance and expected result without including tokens in reports. A final failure decision is not equivalent to waiting for polling. Fixing the cause supports a new evaluation through the authorized process. Also consider time: deferred approval becomes effective only at its selected time and does not resolve other checks. If Business Hours passed but the stage stays blocked until the window closes, time-window approval can be withdrawn. In the scenario, both the external decision and schedule remain unresolved at 23:05. Explain these two blockers to the business and retain the assessed artifact when replanning. Rebuilding for convenience can invalidate previously gathered evidence.
7. Connect capacity, identity and execution limits
A demand helps select an agent; it does not install a missing tool. Confirming actual software, discovering capabilities and exercising the task are different steps. After installing software on a self-hosted agent, restart it during authorized maintenance to refresh discovery. Also distinguish the registration account from the service’s process account: in the Windows example, svc-build reads the share. The administrator’s registration rights do not resolve that operation. For concurrency, maxParallel limits the matrix even with sufficient agents. For duration, a task limit does not extend the job limit. Cleanup eligible after cancellation must also fit cancelTimeoutInMinutes. Ask the learner to calculate remaining time before suggesting more retries. Finally, a ManualValidation wait can use an agentless job, while the following Bash rehearsal needs its agent context and dependency. These distinctions identify the owner who can actually address each failure.
8. Rehearse decisions and communicate what is missing
The code below is a local model of the fictional contract, without Azure calls or YAML interpretation. It covers fourteen boundaries and one hundred result combinations. First predict the two combinations permitting Package and verify that failure, cancellation or an unreachable parent is not accepted. Then compare minutes 21:59, 22:00 and 23:00 and change the external decision from passed to received. Receipt must not satisfy the contract. The simple window stays within one day; time zones and windows crossing midnight are outside this exercise. The model also does not calculate capacity or execute application recovery. Use the results to write a short decision containing evidence, a limitation and the next owner. At work, add service observation and the relevant runbook. Finish the lesson by distinguishing what passed, what never executed and what still needs human acceptance.
# Original fictional acceptance model. No Azure APIs or YAML expression engine.
# Window uses minutes within one day; overnight windows are outside this exercise.
from itertools import product
def can_package(parent_started, cancelled, tests, docs):
return (parent_started and not cancelled
and tests == 'Succeeded'
and docs in {'Succeeded', 'Skipped'})
def ready_to_start(minute, window, checks):
start, end = window
return (start <= minute < end and bool(checks)
and all(value == 'passed' for value in checks.values))
assert can_package(True, False, 'Succeeded', 'Succeeded')
assert can_package(True, False, 'Succeeded', 'Skipped')
assert not can_package(True, False, 'Skipped', 'Succeeded')
assert not can_package(True, False, 'Failed', 'Skipped')
assert not can_package(True, False, 'Succeeded', 'Failed')
assert not can_package(False, False, 'Succeeded', 'Succeeded')
assert not can_package(True, True, 'Succeeded', 'Succeeded')
assert not can_package(True, False, 'SucceededWithIssues', 'Skipped')
checks = {'environment': 'passed', 'external-decision': 'passed'}
assert not ready_to_start(21 * 60 + 59, (22 * 60, 23 * 60), checks)
assert ready_to_start(22 * 60, (22 * 60, 23 * 60), checks)
assert not ready_to_start(23 * 60, (22 * 60, 23 * 60), checks)
assert not ready_to_start(22 * 60 + 50, (22 * 60, 23 * 60), {**checks, 'external-decision': 'pending'})
assert not ready_to_start(22 * 60 + 50, (22 * 60, 23 * 60), {**checks, 'external-decision': 'received'})
assert not ready_to_start(22 * 60 + 50, (22 * 60, 23 * 60), {})
# Two accepted combinations out of 100 in this deliberately strict contract.
states = ('Succeeded', 'SucceededWithIssues', 'Skipped', 'Failed', 'Canceled')
accepted = {(True, False, 'Succeeded', 'Succeeded'),
(True, False, 'Succeeded', 'Skipped')}
count = 0
for inputs in product((False, True), (False, True), states, states):
assert can_package(*inputs) == (inputs in accepted), inputs
count += 1
print(f'14 boundary checks and {count} policy combinations passed')
Prepare pauses a consumer; Deploy fails; Recover is Skipped and Resume never runs. The learner identifies the unreachable path and decides how to confirm and restore service without blindly repeating a change.
Common pitfalls
Using outputs before they exist; confusing YAML position with dependency; interpreting always as an execution guarantee; treating HTTP 200 as a final decision; changing artifacts during approval; assuming a longer task timeout extends the job.
Related topics: Artifact identity · Change management · Incident response · Agents and permissions
Locate the phase and missing condition. A release decision needs evidence about the combination, resources and window; recovery also needs the actual service state.
Reference: Pipeline conditions · AZ-400 objectives 2026-07-27