← AZ-400: DevOps from delivery to operations
17 / 26 · 120 MIN

Tests and coverage: evidence for a release decision

Reconcile execution and results, interpret coverage and instability, and validate representative load before accepting delivery.

1. Define evidence required for the change

In a fictional migration project, release policy requires unit tests, ledger integration, API contract validation and a daily-peak load exercise. Each set answers a different question. A unit test can check a calculation rule without establishing that two applications communicate in the intended configuration. A test using a simulated response can check consumer behavior but does not alone establish that the real provider version honors the contract. Write down each exercise’s boundaries, inputs, artifact version and expected result. Then link evidence to a concrete decision: which risk has been sufficiently assessed and which remains? The expected-suite list must be inspectable. A high count of one test type does not replace a required set that was never executed or whose execution has not been established for the candidate.

2. Separate execution, publication and blocking

The runner executes tests and produces results; the publisher reads files and makes them inspectable; policy decides what blocks progress. A publishing task can succeed with failures in a report if the corresponding control is disabled. It can also tolerate missing files. In this lesson’s case, a directory change leaves integration results outside the search while unit results remain green. Reconciling expected suites with found results avoids confusing absent files with absent failures. Also check format and origin: the publisher’s VSTest option represents TRX and does not require the VSTest runner. That runner’s minimum execution count includes passed, failed and aborted tests; it is not a success count. When many files are published, check tests and configurations inside aggregated runs rather than treating run count as proof of completeness.

3. Produce coverage corresponding to the candidate

Coverage needs collection during a relevant execution, a readable report and linkage to the correct source. PublishCodeCoverageResults@2 publishes previously generated data; enabling failIfCoverageEmpty detects absence but does not execute missing tests. If tests run in a container and the publisher on the host, report paths may not resolve on that host. Check the summary location and use pathToSources when nonabsolute paths need mapping to available source. Do not fill a gap with another commit’s XML merely because modules have the same names. Preserve version identity, parameters and exclusions. Publishing a valid file is also not a quality threshold. The requirement may demand coverage of changes, investigation of critical paths or functional evidence that a percentage alone does not contain. Decide those requirements explicitly before using publication status for release acceptance.

4. Interpret numerator, denominator and paths

Two reports for the same file cover lines 1 through 6 and 5 through 8. Their union contains eight lines; adding six and four duplicates overlap. In distinct modules with 90/100 and 1/10 covered lines, the total is 91/110, about 82.7%, rather than a simple mean of 90% and 10%. The local model lets you check both calculations. Another distinction matters: visiting every line can leave a decision outcome untested. In the exercise’s route function, approved=True visits the lines but does not validate the review result for approved=False. Excluding that path from measurement can improve the percentage without adding evidence. Finally, overall and diff coverage use different populations. In a setup already producing functioning coverage status, an advisory indication needs an appropriate policy if merges below target must be blocked.

5. Address instability without erasing the first failure

A test fails and then passes on rerun without code changes. Retain both attempts: the difference can reveal dependence on state, ordering, timing or environment. In Azure DevOps Services, supported detection can mark the test flaky. That label helps investigation; it does not establish that the cause disappeared. If the team suppresses flaky tests from the summary, the percentage excludes both passing and failing tests in that population. Present the change alongside unreported tests and an owner for correction. Do not compare the new percentage with the old one as if scope were identical. Marking or unmarking affects evaluation in future executions rather than recalculating the current run. A temporary exception should keep unreliable behavior visible, identify required alternative evidence and carry the review date agreed by the team.

6. Optimize execution without losing isolation and scope

Two tests sharing a customer account can pass individually and fail in parallel when one teardown deletes another test’s data. In this scenario, separate data and limit cleanup to resources created by each test. Increasing reruns or inserting a fixed wait does not guarantee isolation. During diagnosis, inspect applied configuration: VSTest task runInParallel can override MaxCpuCount in runsettings. Serial execution can support investigation but does not replace fixing shared state when reliable parallelism is the goal. Impact-based selection also needs boundaries. In a supported TIA environment, changes the analysis cannot understand may trigger all-test execution. Removing that fallback merely to save time replaces a coverage strategy with an assumption of no impact. Document where an optimization is valid and retain evidence of the scope actually exercised.

7. Exercise load and evaluate failure conditions

A completed load test may have no configured criteria. In Azure Load Testing, Completed with No test criteria differs from Passed with evaluated criteria. Define the failure condition in the correct direction: for a fictional minimum of 500 requests/s, failing below 500 is not the same as failing above 500. If the requirement concerns SubmitPayment, apply the criterion to the corresponding request; thousands of fast reads can dilute that operation in the aggregate. High-error auto stop protects execution but does not replace performance criteria. Also inspect the engines generating load: saturated generator CPU can limit throughput before the application. For version comparisons, retain comparable load, data, capacity and scope. Improvement observed with fewer users, more caching and more instances cannot attribute the difference solely to code.

8. Bring sufficient evidence to the committee and RUN

Run the Python model with fictional data. Its teaching policy is deliberately strict: named tests, the correct commit and a first-attempt pass are required. This is neither a universal rule nor a reproduction of the Azure engine. Change an ID, remove a row and introduce a failed attempt before a pass to observe differences an aggregate color would hide. The model also compares line sets and execution counts. In a real handover, supply expected suites, references, configuration, results, attempts, limitations and evaluated criteria. For server metrics in load tests, respect the supported configuration interface and the identity’s metric access rather than assuming any YAML key is effective. If required evidence is missing, recover it or state the pending decision explicitly. The summary is to separate evidence presence, completeness and meaning before approving release.

# Original fictional acceptance model, not an Azure task emulator or a JUnit parser.
# All data is synthetic. The strict policy deliberately accepts only complete first-pass evidence.
from fractions import Fraction

expected = {'unit:pricing', 'integration:ledger', 'contract:payments'}

def acceptable(rows, commit, required):
 keys = [row.get('id') for row in rows]
 if not required or set(keys)!= required or len(keys)!= len(set(keys)):
 return False
 return all(row.get('commit') == commit and row.get('outcomes') == ['passed']
 for row in rows)

rows = [dict(id=name, commit='commit-A', outcomes=['passed']) for name in sorted(expected)]
assert acceptable(rows, 'commit-A', expected)
assert not acceptable([], 'commit-A', expected)
assert not acceptable(rows[:-1], 'commit-A', expected)
assert not acceptable(rows + [rows[0]], 'commit-A', expected)
assert not acceptable(rows, 'commit-B', expected)
assert not acceptable([{**r, 'outcomes': ['skipped']} for r in rows], 'commit-A', expected)
assert not acceptable([{**r, 'outcomes': ['failed']} for r in rows], 'commit-A', expected)
assert not acceptable([{**r, 'outcomes': ['failed', 'passed']} for r in rows], 'commit-A', expected)
assert not acceptable([{**r, 'outcomes': []} for r in rows], 'commit-A', expected)
assert not acceptable(rows + [dict(id='unexpected', commit='commit-A', outcomes=['passed'])], 'commit-A', expected)

# Two report partitions cover the same file revision. Merge by line identity, not by summing hits.
source_lines = {('commit-A', 'payments.py', n) for n in range(1, 11)}
report_a = {('commit-A', 'payments.py', n) for n in range(1, 7)}
report_b = {('commit-A', 'payments.py', n) for n in range(5, 9)}
covered = report_a | report_b
assert len(report_a) + len(report_b) == 10
assert len(covered) == 8
assert Fraction(len(covered), len(source_lines)) == Fraction(4, 5)
assert covered <= source_lines

# Disjoint modules: 90/100 and 1/10. A simple average of percentages distorts the population.
weighted = Fraction(90 + 1, 100 + 10)
unweighted = (Fraction(90, 100) + Fraction(1, 10)) / 2
assert weighted == Fraction(91, 110)
assert unweighted == Fraction(1, 2)
assert weighted!= unweighted

# Minimum execution counts are not a claim that all executed tests passed.
passed, failed, aborted, skipped = 7, 2, 1, 5
executed = passed + failed + aborted
assert executed == 10
assert executed >= 10 and failed > 0
assert executed!= passed + failed + aborted + skipped
print('20 fictional test-evidence checks passed')
IN PRACTICE

A report with 300 passing unit tests lacks required integration evidence because the publisher searches an old path. The case requires recovering candidate evidence and correcting gap detection.

Common pitfalls

Confusing publication with execution; replacing current evidence with old reports; summing repeated lines; hiding flakiness; attributing different load and capacity conditions to code.

Related topics: Acceptance criteria and release governance · Artifact traceability · Capacity and observability · Operational support transition

Take this idea with you

A defensible decision connects criteria, candidate, scope and complete results. Measure what was actually exercised and expose conditions that remain unproven.

Create account

Reference: PublishTestResults@2 reference · AZ-400 objectives 2026-07-27

Microsoft is a trademark of the Microsoft group of companies. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Microsoft. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.