Monitoring and Observability: measure and investigate
Six lessons, 30 questions, and six cases on metrics, traces, instrumentation, alerts, and production diagnosis.
Objectives and progression
A six-module technical course with fictional APS examples. Learn to interpret signals, construct coherent metrics, correlate traces, control cardinality and sensitive data, define alerts, and prepare operational handover. Includes selected OpenTelemetry, Prometheus, Google SRE, and Dynatrace Classic concepts with an internal assessment of 24 decisions in 60 minutes.
Audience: APS L2/L3, infrastructure, SRE, systems administration teams, and technical managers.
Prerequisites: HTTP, application, and production-support fundamentals; examples state scope and limitations.
300 estimated study minutes
- Relate technical signals to useful user operations.
- Interpret rates, percentiles, missing data, and cardinality.
- Follow distributed operations without mistaking partial telemetry for a complete history.
- Prepare useful, controlled, sustainable production telemetry.
- Connect alerts to impact, error budgets, and operational action.
- Combine evidence, synthetic checks, and acceptance criteria for operations.
Modules
- Signals and service outcomes
- Metrics and queries
- Traces and context
- Instrumentation and protection
- Alerts and objectives
- Diagnosis and operational handover
Continue learning
References and version
Observability 2026-09; selected OpenTelemetry, Prometheus and Dynatrace Classic concepts
- Telemetry signals · 2026-09-30
- Spans traces and links · 2026-09-30
- Context propagation and correlation · 2026-09-30
- Head and tail sampling · 2026-09-30
- Sensitive telemetry minimization and redaction · 2026-09-30
- Baggage propagation and trust limits · 2026-09-30
- Collector internal pipeline telemetry · 2026-09-30
- Classic native histograms and summaries · 2026-09-30
- Metric units and label cardinality · 2026-09-30
- PromQL rate aggregation and absent-series functions · 2026-09-30
- Pending firing and alert annotations · 2026-09-30
- Grouping inhibition and silencing · 2026-09-30
- Batch and pipeline instrumentation · 2026-09-30
- User symptoms golden signals and actionable monitoring · 2026-09-30
- SLO burn-rate alerting and traffic limitations · 2026-09-30
- Dynatrace Classic HTTP and browser synthetic monitoring · 2026-09-30
What you will explore
0 / 6Signals and service outcomes
Relate technical signals to useful user operations.
Metrics and queries
Interpret rates, percentiles, missing data, and cardinality.
Traces and context
Follow distributed operations without mistaking partial telemetry for a complete history.
Instrumentation and protection
Prepare useful, controlled, sustainable production telemetry.
Alerts and objectives
Connect alerts to impact, error budgets, and operational action.
Diagnosis and operational handover
Combine evidence, synthetic checks, and acceptance criteria for operations.