← Production Support L3: investigate and recover
11 / 12 · 50 MIN

Observability and distributed evidence

Interpret traces, metrics, and problems without confusing observation gaps with business outcomes.

Connecting evidence across services

A valuation request may cross an API, worker, and data service. Start investigation with the operational identifier, time window, and known components. Tracing context helps connect spans across processes when propagated correctly. In an exercise, logs establish worker execution but the trace ends at the API: inspect the HTTP boundary and instrumentation before concluding that the worker received no work. Keep the business-operation ID distinct from the trace ID. Their relationship should be observable without exposing unnecessary financial content. Record which boundary has evidence and which still needs investigation.

Sampling and missing traces

Trace search should not be treated as a ledger. Sampling retains some observations, and a policy favoring errors changes the visible failure proportion. In a fictional case, 40% of stored traces contain errors because the policy favors them; that value is not automatically the overall service failure rate. For a disputed instruction, inspect operating state and telemetry coverage. Record the window, filters, and ingestion delays as well. A useful conclusion may be “execution confirmed, trace unavailable,” leaving the observability gap as a separate follow-up action.

Counters, resets, and missing data

Before interpreting a chart, identify metric type and query interval. A counter can reset when its process restarts; the drop does not represent negative requests. In Prometheus, calculate rate per series before summing instances to preserve reset handling. In an exercise, one process restarts while another receives more traffic: a raw sum can hide the reset. If scraping failed, “no data” is not zero either. Cross-check ingress or queue counts and state the age of the last valid sample in the operational update.

Percentiles and comparable populations

A p95 describes a distribution of observations. Adding or averaging two servers’ p95 values does not recover the percentile of their combined requests, even with volume weighting. Use aggregatable distributions such as compatible histograms when a global result is needed. For two clusters, also retain analysis of a critical flow passing through only one. Even a correct global statistic can hide an important segment. Tell the manager which populations were observed, the window, and agreed criteria; do not fill a data gap with invented precision.

Labels that support decisions

Choose metric dimensions that answer operating questions with a bounded set of values. Environment, service, and route template support segment comparison. A full URL containing IDs or one label per trace can create a growing number of series. In a design exercise, failures must be located by endpoint: use /orders/{id} as a normalized route and retain the individual identifier in appropriate access-controlled evidence. Avoid collecting tokens or account numbers for convenience. Review each dimension’s cost, usefulness, and necessity before expanding instrumentation.

Dynatrace problems and functional acceptance

In Dynatrace, a problem relates events and dependency context to guide investigation. Inspect affected entities, evidence, and impact rather than relying only on the first alert received. In a fictional incident, the latency problem closes after dependency recovery, but 300 instructions remain unreconciled. The update should state technical recovery and outstanding business work, with an owner and next communication time. The tool lifecycle does not replace service acceptance criteria. Always summarize what telemetry supports and what requires another form of evidence before handing work over.

IN PRACTICE

An instruction is confirmed in the ledger but absent from sampled tracing. Reconcile the operational identifier before authorizing repetition.

Common pitfalls

Sample as full population; averaged percentiles; missing data as zero; alert closure as acceptance.

Related topics: Shift handover and escalation · Runbooks that support decisions · Linux, JVM, and connectivity diagnosis

Take this idea with you

Every conclusion should state population, window, coverage, and independent outcome evidence.

Create account

Reference: Context propagation · DR Production Support L3 2026.4; Linux, JDK 25 HotSpot, OpenSSL 3.5 and Kubernetes examples require installed-version checks