← AZ-400: DevOps from delivery to operations
20 / 26 · 90 MIN

Observability: coverage, retention and freshness

Establish that signals represent the components and time of the decision, from collection through diagnosis and RUN handover.

1. Start from the operational decision

A dashboard helps only when its signals answer the operational question. In a fictional funds service, APS needs to distinguish infrastructure saturation, dependency failure and delay in telemetry itself. Define which components should emit data, how often and to which destination. For each chart or alert, record unit, aggregation, dimensions, time interval and owner. An old CPU value does not establish stability during a new incident. An overall average does not establish every instance is healthy either. Start handover with these observable conditions and plan a controlled exercise. Dashboards with the right names constitute a configuration inventory; operational acceptance requires demonstrating that current data arrives and someone can interpret it.

2. Confirm the machine collection chain

In machine monitoring with Azure Monitor Agent, distinguish installation, rule and association. The DCR defines sources and destinations; association links that rule to the machine. A correct rule that was never associated does not establish collection. After migration, confirm resource identity, effective configuration and the workspace used by the query. Current documentation distinguishes OpenTelemetry metrics and the classic logs-based experience, so do not assume one table fits every monitoring path. One hypothesis in the lesson case is that the dashboard still queries the old destination and shows earlier data. Investigation checks that hypothesis instead of immediately declaring an application or agent failure. Continue incident response in parallel, using independent signals when current and suitable.

3. Align streams and scope in Kubernetes

In AKS, different data follows different collection paths. When moving ContainerLogV2 to high-scale mode, review mode and stream configuration. Retaining both the normal and high-scale streams can duplicate records. This affects cost and interpretation, not merely presentation. Do not solve the problem by dividing every count by two: first establish which sets were duplicated, over what period and under which configurations. Similarly, namespace filters do not affect every table or metric uniformly; a node is not a namespace-scoped object. Give support the list of sources and exclusions that actually applies. Expected absence caused by filtering and unexpected absence caused by collection failure require different responses.

4. Assess what filtering loses

Reducing volume can be a legitimate cost decision, but it should start from events’ operational value. An ingestion transformation acts before workspace storage. If it removes an event class, increasing retention later does not recover events never retained unless a recoverable source or copy exists. In the exercise, the source has rotated the records and no other copy exists; investigation must acknowledge the gap. Before changing filters, try them against a representative sample and identify diagnostic questions that would become unanswerable. Record change date, owner, scope and expected cost effect. Keep the decision tied to risk: removing repetitive noise is not automatically equivalent to removing events needed to reconstruct an incident.

5. Distinguish event time and observation time

A sample can arrive late without having been generated late. TimeGenerated represents the creation time supplied by the source, while receipt and ingestion describe later stages. In the example, 10:00, 10:07 and 10:08 separate seven minutes to receipt from one more minute to ingestion. Do not call those eight minutes application transaction duration: transaction timestamps are missing. Also check whether timestamps were adjusted and which column filters the query. If the dashboard uses event time, newly ingested data can appear in an earlier window. When investigating delay, retain clock context, resource and analyzed interval. A time difference indicates where to investigate; it does not alone establish network, agent or processing as the sole cause. When teams work in different time zones, record the agreed time basis in the incident timeline so each comparison uses the same instants.

6. Query history according to plan and retention

Analytics and total retention are not periods to add together. In an Analytics table with 30 analytics days and 180 total days, retained records from 90 days ago may require retrieval through a search job rather than the usual query. First establish that records were ingested and remain retained, along with the table plan and configuration. Do not generalize the same mechanism to every plan. This distinction matters in post-migration investigation: not finding a row in one view does not establish deletion, but neither does it justify promising that the row exists elsewhere. Define acceptable historical-retrieval time and its operational cost in advance. Translate the audit or diagnostic requirement into a rehearsed procedure and owner, not merely a day count.

7. Preserve metric identity and population

Application map uses cloud role name to identify components. Two applications with the same value can appear grouped; this does not prove they are the same process or that one is redundant. Maintain stable component identity and use instance information to investigate replica differences. Also preserve calculation populations. With values 80, NULL and NULL, only one valid measurement exists: the average is 80 with incomplete coverage. Replacing absence with zero would fabricate two measurements and produce about 26.7. Do not interpret that artificial drop as improvement. In management reports, show freshness and coverage alongside the value, especially for capacity, migration or incident-closure decisions. Observation quality is part of result interpretation.

8. Exercise handover with explicit data

The records below are an original exercise, not an Azure export. Calculate intervals and average, identify the stale dashboard and propose the next check. The solution should separate known state from remaining hypotheses. The changed DCR may explain lost visibility, but it is not established as the cause of service errors. Then plan a controlled alert and follow notification through to the owner and diagnostic procedure. Record source, destination, dimensions, frequency, queries and limitations. If critical signals are missing, keep acceptance pending or use an explicitly validated alternative. The RUN summary should explain what to do when old data arrives, when data is absent and when current data shows a real failure.

{
 "fictional": true,
 "record": {
 "TimeGenerated": "2026-10-05T10:00:00Z",
 "TimeReceived": "2026-10-05T10:07:00Z",
 "IngestionTime": "2026-10-05T10:08:00Z",
 "timestampAdjusted": false
 },
 "metricPositions": [
 80,
 null,
 null
 ],
 "dashboard": {
 "lastSample": "2026-10-05T10:00:00Z",
 "now": "2026-10-05T10:20:00Z",
 "expectedFreshnessMinutes": 5
 }
}
IN PRACTICE

After migration, a chart retains a twenty-minute-old sample while the service reports errors. The case requires restoring visibility and continuing incident investigation.

Common pitfalls

Confusing agent installation with complete collection; duplicating streams; trying to recover discarded data with retention alone; averaging invented zeros; treating a map as proven architecture.

Related topics: DCR and configuration management · Incident diagnosis and data freshness · Retention and historical retrieval · Alerts and operational handover

Take this idea with you

Check identity, coverage and time before interpreting a value. A useful operational signal reaches the correct destination, represents the relevant time and enables a demonstrable response.

Create account

Reference: Enable VM monitoring in Azure Monitor · AZ-400 objectives 2026-07-27

Microsoft is a trademark of the Microsoft group of companies. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Microsoft. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.