← AZ-400: DevOps from delivery to operations
26 / 26 · 100 MIN

Operational KQL: coverage, correlation and decisions

Analyze telemetry without losing legitimate events, hiding missing sources or confusing estimates with exact counts.

Start with the population the decision requires

Before writing the query, define what must be represented: services, environments, versions and observation interval. In an API and Batch release exercise, one healthy API row does not represent the whole set. Keep an independent list of expected deployments and compare it with received telemetry. ExpectedDeployments on the left of a leftanti against Heartbeats identifies deployments without a match in the chosen window. The result is a list of evidence gaps, not an automatic failure diagnosis. For each gap, identify the source owner, confirm destination and filters, and record a deadline for clarification. This preparation prevents the query itself from silently defining the population it later declares healthy. The decision should explain both observed services and those whose coverage remains unconfirmed.

Control cardinality before correlating

Two events can share OperationId without being duplicates. A payment has accepted and settled events; both are needed to analyze its journey. Default join uses innerunique and removes repeated left keys, so one event may be lost. Do not promise which survives. If the owner table has exactly one matching row, join kind=inner retains both. If it has two, correlation can multiply results. Therefore write cardinality expectations for each side and compare counts before and after. Sorting input or changing an execution hint does not correct the wrong semantics. Nor should inner measure deployment coverage: unmatched deployments disappear. Choose the operation for the question and separately retain evidence of unmatched items. A smaller result is not automatically a cleaner or more accurate result.

Preserve the state of the latest row

Deployment panels often need the latest state and when it was observed. max(TimeGenerated) calculates a time value but does not itself select the corresponding State. Combining that maximum with an arbitrary state can create a row that never existed in the source. Use summarize arg_max(TimeGenerated, State) by DeploymentId when the desired state belongs to the row maximizing time. Exercise timestamps are unique per deployment, avoiding tie ambiguity. In real data, define a rule for different events sharing a timestamp and confirm the time field’s meaning. A recent state still needs an environment and candidate identity. In the RUN report, present that full identity and the selected window so nobody confuses the latest test event with production status or reads an old observation as current.

Interpret the type before interpreting the value

A producer may send a JSON object whose detail field is another JSON string. After parse_json(Payload), that field is dynamic but still contains text. Calling parse_json directly on the dynamic field again passes it through unchanged. The expression parse_json(tostring(d.detail)).code first converts the field to text and then parses the inner JSON. This detail appears during log investigations when a column seems to contain information but extraction returns an unexpected value. Before changing alerts, inspect an authorized sample and confirm the producer’s contract. Distinguish malformed JSON, an absent property and an internal string not yet parsed. The exercise assumes valid JSON in both layers; it neither promises to repair invalid formats nor establishes that all events follow the same contract.

Treat warnings as part of the result

A query can return useful rows and still be incomplete for the decision. Within the same existing database, union isfuzzy=true may resolve ApiLogs, fail to resolve BatchLogs and return data with a warning. Tolerance of partial resolution is not evidence that Batch had no errors. It also does not indiscriminately suppress later query failures. Record the warning alongside expected and actually observed sources. If both are mandatory, promotion depends on recovering missing evidence. An alternative source only works if approved for the criterion and matched to the same candidate and interval. Avoid implicit source lists that change scope without review. A project team should be able to explain which dependency remains open, who resolves it and how it affects the delivery window.

Separate empty intervals from zero activity

bin(TimeGenerated, 5m) groups events into intervals; it does not create rows for periods without observations. Adding buckets to the axis can improve readability, but a presentation zero does not establish complete collection. When the collector is unconfirmed, identify the gap. Boundary choice also matters: between includes both endpoints. Two consecutive reports can count the same event at their meeting point. Using TimeGenerated >= Start and TimeGenerated < End assigns that event to the next window while including each start. This convention resolves temporal overlap, not source duplicates or late arrivals. In a handover, record timezone, boundaries, extraction time and coverage state. Differences between reports can then be investigated without immediately assuming changed service behavior or silently adjusting the reporting population.

Choose between estimation and exact counting

For exploration across millions of identifiers, dcount offers an estimate with accuracy and resource tradeoffs. Exact modes for small sets do not guarantee exactness at every size. If the requirement is an exact CustomerId count from a fixed extract, first define the population, empty-value handling and query completeness. In the exercise, all IDs are nonempty and execution fits approved limits; distinct CustomerId followed by count counts the observed set. Do not replace this process with a rounded mean of repeated estimates. Nor should events be confused with customers: one customer can have many rows. Finally, an exact calculation over an incomplete extract still does not represent all real customers. Communicate scope alongside the number and retain the criteria used to produce the extract.

Guided exercise and release decision

In the example below, the expected list contains D1 and D2, but only D1 has a heartbeat. Predict the query result before running it in an authorized learning environment: D2 should appear as an evidence gap. This static example has not been executed in a Kusto engine. Then analyze the final case: API is green, BatchLogs was unresolved and policy requires both services. Write a note holding promotion, naming the investigation owner and specifying evidence needed for reconsideration. Do not declare Batch broken merely because logs are missing. A current approved alternative can close the gap; an old result cannot satisfy the criterion. Finish by comparing expected population, observed rows, warnings and time boundaries before interpreting any dashboard color as delivery authorization.

// Static fictional exercise; not executed in a Kusto engine.
let ExpectedDeployments = datatable(DeploymentId:string) ["D1", "D2"];
let ObservedHeartbeats = datatable(DeploymentId:string) ["D1"];
ExpectedDeployments
| join kind=leftanti ObservedHeartbeats on DeploymentId
// Predicted result: D2. Missing evidence, not a confirmed service failure.
IN PRACTICE

API and Batch are required for release acceptance. The query shows only API and warns about BatchLogs: the result does not establish sufficient coverage.

Common pitfalls

Deduplicating legitimate events, hiding deployments in an inner join, filling gaps as success and assigning universal exactness to dcount.

Related topics: Release acceptance criteria · Telemetry quality and contracts · Incident investigation and handover

Take this idea with you

A useful query preserves the population required by the decision and distinguishes result, coverage, precision and uncertainty.

Create account

Reference: innerunique join · AZ-400 objectives 2026-07-27

Microsoft is a trademark of the Microsoft group of companies. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Microsoft. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.