Concept and mechanism
An observability platform should help explain service behavior. Metrics aggregate measurements over time; logs record events; traces connect operations across a distributed path. Each signal has limits. Normal CPU does not establish that a transfer was accepted, and one log entry does not automatically represent every request. OpenTelemetry provides mechanisms to instrument, collect, process, and export telemetry; using it does not remove the need to choose suitable storage, queries, and visualization. Start by defining a relevant operation, its expected outcome, and the measured population. In a fictional position-query service, distinguish valid requests, technical failures, and expected business-rule rejections.
Guided application
Use latency, traffic, errors, and saturation for an initial assessment. If error responses are very fast, overall average latency can improve during a failure; separate outcomes and examine the distribution tail. Calculate ratios using explicit denominators: 60 errors among 12,000 eligible requests represent 0.5%. Do not take a simple average of instance percentages when traffic volumes differ. Combine the external perspective, which exercises the service, with internal signals that help locate causes. For a batch workload, a live process is insufficient: establish when it last completed successfully and whether expected data was produced. Record the window, unit, and meaning of each signal on the dashboard.
Lower average latency with more fast HTTP 500 responses can indicate degradation.
Common pitfalls
CPU as experience; mean as tail; expected rejection as technical failure; percentages without denominators.
Related topics: Metrics and queries · Traces and context · Instrumentation and protection
Define the useful outcome before choosing the indicator.
Reference: User symptoms golden signals and actionable monitoring · Observability 2026-09; selected OpenTelemetry, Prometheus and Dynatrace Classic concepts