1. Start with the question and population
The PM of a fictional release receives two counts: A had ten failures and B had four. Before reporting improvement, ask how many requests each version received, in which window and for which operation. In the exercise, A has one thousand requests and B one hundred, with complete recording and no sampling. Observed rates are one and four percent. The lower absolute number does not establish improvement. This also does not prove version causality: load, customers and operations may need further analysis. Define the population the decision should represent and retain numerator and denominator in the report. For a canary, show exposure and outcomes by version instead of mixing everything into one global average. Include collection time and failure-classification criteria. Another person should be able to reproduce the comparison from identified data.
2. Count rows or represented items
AppRequests includes ItemCount, the number of items represented by a sampled row. In the controlled example, the failure row represents ten requests and the success row ninety. Counting one failure among two rows gives fifty percent of rows; using weights gives ten percent of represented requests. Explain which question each calculation answers. In KQL, group by the intended version and calculate sums and conditional sums when weights apply. Confirm schema and types because workspace tables do not necessarily use the same names as earlier query experiences. Do not invent weights for data lacking that information or use arithmetic to hide exporter losses. The exercise assumes weights correctly represent the population; real operation requires validating that assumption. An empty population should appear as missing evidence for assessment rather than zero percent failures.
3. Do not aggregate summaries as if they were raw data
Two instances report p95 values of one hundred and nine hundred milliseconds. Their mean, five hundred, is not automatically global p95. In the laboratory, we create two populations with those same individual summaries: in one, the fast instance receives almost all traffic; in the other, the slow instance receives almost all traffic. Global percentiles differ. The example uses exact nearest-rank over small local lists; it does not reproduce Kusto’s approximate algorithm. For real analysis, retain data or distributions suitable for intended aggregation and interpret the chosen method’s limitations. Even a weighted mean of the two p95 values does not generally reconstruct the lost distribution. During reporting, identify whether a number represents an instance, operation, region or combined population. A dashboard can display well-formatted numbers answering different questions. Confirm meaning before using chart color as a promotion criterion.
4. Reconstruct the path without inventing causality
In the tracing example, A and B share TraceId and B’s ParentSpanId points to A’s SpanId. This identifies B as A’s child within the observed path. Do not delete B as a supposed duplicate merely because both share the trace. Context propagation links execution units but does not establish that a business outcome was accepted at the destination. When investigating a payments timeout, relate the span to appropriate functional confirmation and distinguish submission, receipt and outcome. If some context is not propagated, temporal proximity alone cannot justify inventing a parent-child relationship. Document the gap and inspect service boundaries. Use diagnostic correlation identifiers without placing personal data or secrets in fields. Handover should explain which path was observed and where evidence is still missing. Tracing helps structure investigation rather than replace business reconciliation.
5. Distinguish alert state from repeated messages
A stateful rule remains Fired while its condition is unresolved and does not fire again on every evaluation. The team should track the active alert and its owner rather than expect repeated messages to confirm persistence. Define an operational update and escalation policy that does not depend on that assumption. Another gap appears on a new resource whose dynamic thresholds lack sufficient history. Absence of an alert during learning is not validation of the first release. Consult current requirements for that rule type and retain explicit observation and acceptance criteria during the period. Avoid choosing a day count from memory or extrapolating support between metrics. For each signal, record condition, population, window, state, intended action and owner. That description helps identify whether detection, interpretation or human response failed and which team should address it.
6. Rehearse the path to the responder
In the fictional case, the batch failed at 22:55 during maintenance. The alert exists in the portal but an alert processing rule suppressed action groups. At 23:10, the deadline is approaching and nobody owns response. Window end does not guarantee retroactive execution of suppressed actions. Invoke the authorized operational path and confirm batch state while reviewing rule scope, filters and schedule. A second rule adding action groups does not override applicable suppression. Do not restart business work merely to produce a fresh test alert. Separate notification rehearsal from application recovery and confirm partial outcomes before any repetition. During handover, identify who watches alerts during maintenance, how support is contacted and who confirms exit from suppression. Completing the change includes restoring that response capability. Record both detection evidence and acknowledgement by the accountable responder.
7. Connect probes to functional criteria
The exercise’s /health endpoint only confirms that the web process responds. HTTP 200 does not establish completion of a batch the endpoint never queries. Adding locations or frequency to the same probe improves observation of that test but does not add the missing functional path. Define suitable batch signals: expected start, progress, completion and the outcome business owners need confirmed. Concrete criteria belong to the application and should be agreed with accountable people. A synthetic test also needs to reflect what it claims to validate; naming a probe Payments does not automatically give it payments coverage. During a release, compare technical signals with functional outcomes without equating them. If the PM received only web evidence, communicate that limitation and assign batch confirmation. This avoids closing a change with an important part of service still unobserved.
8. Practise calculations and write an auditable decision
Execute the local code after predicting both rates and combined percentiles. Its fourteen checks use fictional data and neither execute KQL, query Azure nor measure real telemetry quality. The small alert model is a snapshot of supplied conditions, not a scheduler or message-delivery simulator. Change suppressed while retaining fired as true and observe that visibility and notification differ. Then change owner and identify when response ownership is still missing. Finish by writing four lines: observed population, result, limitation and next action with an owner. In the batch case, explicitly state the lack of guaranteed replay and the need to confirm outcomes before restarting. Connect this lesson to release recovery and handover: a technically correct decision takes effect only when it reaches someone able to act and is recorded with sufficient supporting evidence.
# Original fictional datasets; local arithmetic only, not a KQL execution engine.
# Weighted rows are assumed representative. No claims about sampling correction
# beyond the supplied ItemCount values or production telemetry completeness.
from math import ceil
def weighted_failure(rows):
total = sum(r['ItemCount'] for r in rows)
return None if total == 0 else sum(r['ItemCount'] for r in rows if not r['Success']) / total
def nearest_rank(values, percentile):
return sorted(values)[ceil(len(values) * percentile / 100) - 1]
assert 10 / 1000 == 0.01
assert 4 / 100 == 0.04
assert 4 < 10 and 4 / 100 > 10 / 1000
sample = [{'Success': False, 'ItemCount': 10}, {'Success': True, 'ItemCount': 90}]
assert weighted_failure(sample) == 0.1
assert sum(not r['Success'] for r in sample) / len(sample) == 0.5
assert weighted_failure([]) is None
# Two datasets have identical per-instance p95 values but different global p95.
large_fast = [100] * 1000
small_slow = [900] * 10
small_fast = [100] * 10
large_slow = [900] * 1000
assert nearest_rank(large_fast, 95) == nearest_rank(small_fast, 95) == 100
assert nearest_rank(large_slow, 95) == nearest_rank(small_slow, 95) == 900
assert nearest_rank(large_fast + small_slow, 95) == 100
assert nearest_rank(small_fast + large_slow, 95) == 900
assert (100 + 900) / 2 == 500
def decision(fired, suppressed, owner):
# Snapshot model: no Azure scheduler, rule propagation or message delivery.
return {'visible': fired, 'notify': fired and not suppressed,
'needs_owner': fired and not owner}
assert decision(True, True, False) == {'visible': True, 'notify': False, 'needs_owner': True}
assert decision(True, False, True) == {'visible': True, 'notify': True, 'needs_owner': False}
assert decision(False, False, False) == {'visible': False, 'notify': False, 'needs_owner': False}
print('14 arithmetic and decision assertions passed; no KQL or Azure alert execution')
The canary version has fewer failures but a higher rate; a batch alert fired during suppression and remained unowned. The learner calculates populations and decides how to initiate response without a blind restart.
Common pitfalls
Comparing counts without exposure; ignoring ItemCount; averaging p95 values; confusing tracing with success; expecting repeated stateful-alert messages; assuming post-maintenance replay; treating a web probe as batch validation.
Related topics: KQL and aggregations · Distributed tracing · Incident management · RUN handover
Define each signal’s population and limitations. Confirm detection, routing and response ownership before declaring operational coverage established.
Reference: Use aggregation functions in Kusto Query Language · AZ-400 objectives 2026-07-27