Concept and mechanism
Metrics, logs, and traces answer different, complementary questions. A metric summarizes behavior over time; a log describes an event; a trace links request operations when instrumentation and propagated context exist. For a slow valuation crossing an API, queue, and calculation stage, global average CPU does not necessarily reveal where time was spent. Define the diagnostic question before choosing a chart. Also consider what was not collected: a missing trace does not prove the request never occurred. Observability design should relate application and dependencies without depending on unrestricted access to sensitive data.
Guided application
A service objective needs an indicator, population, and window. In a request-based example, 99.9% success among one million eligible requests permits 1000 failures; with 900 already counted, 100 remain in that same fixed population. Do not convert that result into minutes without defining a time-based indicator. Tool selection also requires context. CNCF maturity informs the project’s sustainability but does not automatically provide an SLA for your installation. Assess requirements, licensing, upgrades, skills, and operational ownership. In an international team, document decisions and acceptance criteria so development, platform, and APS can work from the same evidence.
At production handover, present the observed transaction, objective, alert recipient, and owner of error-budget decisions.
Common pitfalls
CPU treated as every request’s latency; missing trace treated as no event; percentage without denominator; graduated treated as an SLA.
Related topics: Desired state, control, and containers · Capacity, scheduling, and workload types
Signals and tools help when connected to an objective and concrete owners.
Reference: Telemetry signals · KCNA current four-domain curriculum; edition date unconfirmed