← AZ-305: Azure architecture and production decisions
16 / 23 · 100 MIN

Observability: collection, alerts and cost

Design an evidence path from source to operational decision with explicit retention, access and cost.

1. Start with the question support needs to answer

In a fictional funds application, identifying who changed configuration differs from explaining a failed data read. Activity Log records management operations; resource logs and application telemetry provide other perspectives. Do not ask support to find every failure in one record type. Define concrete questions: was the change applied, which dependency denied access, which business operation remains outstanding? Link each question to a source, correlation identifier and destination. A successful deployment entry does not demonstrate that the application processed its first instruction. Include what is not collected by default and how a representative event will be generated to validate collection before handover. This makes the observability requirement a diagnosable workflow rather than simply a list of enabled services.

2. Distinguish created resources from working collection

A created workspace and installed agent do not prove required records are arriving. Diagnostic settings select resource-log categories and destinations. For Azure Monitor Agent collection, data collection rules and relevant associations define what to collect and where it goes. In the fictional example, the host emits platform metrics but no operating-system events appear: inspect the agent, rule, association and destination path rather than conclude that missing visible errors mean a healthy VM. Some tables appear only after first ingestion. Acceptance should include the generated event, observed record and measured delay. An empty dashboard can mean no activity, ingestion delay or collection failure. Assign an owner to distinguish those states before an incident occurs.

3. Transform data without losing needed evidence

An ingestion transformation can filter or reshape data on supported paths and tables. It can remove unnecessary fields, but a filter discarding all successes can also remove the denominator of a failure rate. If the team calculates ten errors among one hundred operations, it needs a reliable way to retain the total even when reducing success detail. Do not treat a dashboard query as equivalent to removing data before storage: they act at different pipeline points. Prepare samples containing success, failure, missing fields and fictional identifiers. Compare results before and after transformation, confirm diagnostic usefulness and verify that required cross-component correlation remains. Do not approve a transformation solely because it reduces volume; the operational decisions using those records still need defensible inputs.

4. Choose plan and retention from usage

A table plan affects available capabilities and ingestion and query cost. Data continuously used by alerts differs from high-volume records consulted only during occasional investigations. Before moving from Analytics to a cheaper plan, inventory queries, alerts, summary rules and users. Current documentation identifies specific limitations, including alerts stopping after a move to Auxiliary or Lake. Do not treat a plan change as automatic deletion of earlier data either: confirm access methods for each period and applied retention. In an international project, separate data residency from dashboard organization. A dashboard per country does not change the storage location of a shared workspace. The choice should preserve required investigations and detection paths while making any operational trade-off visible to their owners.

5. Separate detection, notification and response

An alert rule detects a condition; an alert processing rule can change actions applied to already-fired alerts. During maintenance, suppressing notifications for a bounded scope and schedule can be appropriate while retaining evidence of observed conditions. Disabling detection has a different effect and can hide failures on resources outside maintenance when they share the same rule. Define the window owner, time zone, resources and how notification resumption will be confirmed. Receiving an email also does not prove someone assumed incident ownership. Handover should connect signal, severity, notification destination, responsibility and response procedure. The local model distinguishes missing data, no traffic, a detected condition and suppressed notification. It does not execute Azure alert logic or guarantee real notification delivery.

6. Avoid savings that create blind periods

A daily cap can stop eligible ingestion after a threshold, but it is neither an exact billing ceiling nor a consequence-free routine filter. Stopping collection can remove alerting and diagnostic data precisely during an incident surge. Cost modeling should consider covered tables and plans, possible overshoot, resumption and monitoring of the cap itself. In the fictional case, the team lowers the cap and the afternoon dashboard shows zero errors. Before reporting improved availability, confirm collection and operation volume. Plan selective detail reduction using samples and validation, accompanied by warnings before the cap is reached. The intended outcome is controlled cost while preserving agreed operational decisions and evidence. A cheaper dashboard that cannot distinguish silence from health does not meet that requirement.

def assess(total, errors, complete, notify):
 if total < 0 or errors < 0 or errors > total:
 raise ValueError("invalid fictional counts")
 if not complete:
 return ("unknown", False)
 if total == 0:
 return ("no-traffic", False)
 fired = errors / total >= 0.05 # fictional threshold, not an Azure rule
 return ("fired" if fired else "clear", fired and notify)

assert assess(0, 0, False, True) == ("unknown", False)
assert assess(0, 0, True, True) == ("no-traffic", False)
assert assess(100, 10, True, True) == ("fired", True)
assert assess(100, 10, True, False) == ("fired", False)
assert assess(100, 1, True, True) == ("clear", False)
assert assess(100, 10, False, True) == ("unknown", False)
print("six fictional signal checks passed; no Azure alert evaluated")
IN PRACTICE

Fictional case: after reducing ingestion, success records disappear and the failure rate becomes misleading. The owner preserves reliable aggregate counts, reduces selected detail and retests detection and notification.

Common pitfalls

Confusing a workspace with collection; interpreting silence as health; removing metric denominators; confusing notification suppression with no alerts.

Related topics: SLA and operational indicators · FinOps and retention · RUN handover

Take this idea with you

Observability needs usable data, correct detection and an owned response, including during cost changes and maintenance.

Create account

Reference: Diagnostic settings in Azure Monitor · AZ-305 objectives 2026-04-17

Azure is a trademark of the Microsoft group of companies. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Microsoft. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.