← AI-200: cloud development for AI solutions
06 / 6 · 40 MIN

Telemetry, KQL, and diagnosis

Interpret metrics and traces with explicit denominators, windows, and coverage.

Concept and mechanism

OpenTelemetry provides instruments for observing operations and dependencies; Application Insights helps analyze distributed paths. Metrics, logs, and traces complement each other. Sampling reduces volume but changes the available set of individual events: a missing trace does not prove a request never occurred. For indicators, define population, time window, filters, and sampling treatment. In KQL, countif counts rows satisfying a condition; count counts the observed population. If you first filter to failures only, the denominator no longer represents every request. Percentiles are useful estimates for observing latency distribution; the mean can hide a slow tail. Compare equivalent segments by operation, version, and time window.

Guided application

In a complete fictional set with 240 requests and 12 failures, the rate is 100 × 12 / 240 = 5%. The calculation assumes no sampling and matching numerator and denominator scope. An illustrative query can aggregate total=count and failures=countif(not(success)), then calculate the ratio; adapt table, types, and fields to the actual workspace. This course did not execute KQL against a service. For latency, if the database consumes eight of nine seconds, investigate queries and waits before increasing API replicas. Retain correlation and minimal data. Baggage can cross service boundaries, so tokens and personal data should not become diagnostic shortcuts. Communicate evidence limits when reporting is inconclusive.

IN PRACTICE

Twelve failures in 240 requests = 5%. Filtering to the 12 failures before counting total would produce an incorrect denominator.

Common pitfalls

No trace as zero failures; span count as request count; mean as the experience of every user; percentile as an exact measurement; baggage as a vault.

Related topics: Containers, artifacts, and revisions · Cosmos DB, vectors, and change feed · PostgreSQL, vectors, and caching

Take this idea with you

Explain what was observed and what coverage allows you to conclude.

Create account

Reference: Distributed application observability · AI-200 current guide updated2026-05-05; Azure technical documentation2026-09-30