Define population and denominator before the percentage
A percentage is interpretable only with population, interval and counting rule. In the exercise, A has 9 failures in 9000 requests and B has 20 in 1000, over the same window without overlap. Overall failure rate is 29/10000=0.29%. The simple average of 0.1% and 2% would be 1.05%, because it gives equal weight to groups with different volumes. A correct overall rate can nevertheless hide degradation in a small population. Retain dimensions needed for decisions, such as revision or business journey, alongside totals. Before comparing dashboards, check whether they count requests, attempts or completed operations and whether exclusions match. In an APS handover, deliver these definitions with queries and owners. An apparently precise number without those definitions can lead development, production and business teams to incompatible decisions. Check the observation window before attributing a change to a release.
Recognize the precision available in histograms
An aggregated distribution stores counts per interval; it does not necessarily retain each original latency. In the example,90 measurements lie in [0,100) ms and 10 in [100,1000) ms. Defining observed p95 as the 95th of 100 sorted measurements tells us it lies in the second interval. We do not know whether it was 120, 550 or 900 ms. An estimate displaying 550 does not establish an exact observation. To practise, construct two datasets with identical counts, changing only the ten values in the second interval, and compare observed percentiles. Then discuss whether bucket width supports the intended latency decision. Do not add apparent reporting precision through decimal places the data cannot support. Distinguish count, estimate and individual measurement when explaining results to the service approver. Keep the original bucket model alongside the chart so another operator can interpret its limits.
Interpret alert logic and resource identity
Metric-alert alignment and retesting serve different purposes. Alignment regularizes observed points; retesting requires the condition to persist for its configured time. If an already aligned value no longer violates the threshold, the retest window resets. Do not add two high periods separated by a healthy value as though they were continuous. Combining conditions also matters: AND can combine high CPU on one VM with high memory on another. To require both on the same VM, evaluate AND_WITH_MATCHING_RESOURCE and preserve compatible resource labels after aggregation. In the exercise, draw two timelines, one per VM, and mark which condition holds on each. Then explain the operational decision the alert should enable. A dashboard name cannot repair a query that lost required identity; configuration must preserve that relationship. State missing-data behavior separately when defining the actual policy.
Separate storage, counting and trace coverage
A user-defined project-scoped metric can count events received by the Logging API even when an exclusion prevents storage in _Default. The exercise assumes active billing, an existing metric and events matching its filter. Do not generalize to a bucket-scoped metric: that depends on logs stored in its bucket. A rising counter also cannot reconstruct discarded messages. Keep aggregated evidence separate from retained content in the design. In tracing, missing request spans do not invalidate errors established by counters or logs. Sampling and context propagation are distinct mechanisms; a trace ID does not guarantee complete collection. During an incident, retain observed errors and treat missing traces as a diagnostic gap. Check sampling configuration, instrumentation and limits before concluding the request did not exist or a downstream call never occurred. Document what each evidence source can actually establish.
Choose suitable diagnostic signals and probes
Java profiles with high wall time and low CPU time suggest waiting, not an already established single cause. Wall time includes waits for I/O, locks or synchronization; CPU time measures processor execution. Use the difference to choose further evidence, such as stacks and dependency latency. Do not approve more vCPU solely because a function takes a long time. In GKE, legitimate 90-second startup can be interrupted by premature liveness. False readiness does not suspend liveness. A startup probe delays readiness and liveness until it succeeds, with tolerance configured for measured startup. Restart remains possible if startup probing fails beyond tolerance. The exercise asks for two separate criteria: acceptable initialization time and a condition indicating inability to make progress afterward. This avoids turning temporary slowness into restarts that prevent recovery. Measure both behaviors under representative workload and configuration.
Reconcile work after a timeout
A Cloud Run request timeout closes the connection and returns 504, but does not terminate the container serving the request. Code may continue; successful completion is not guaranteed either. Therefore, transport error does not establish absence of effects. In the report-41 case, a durable record becomes completed and references the published file after timeout. Consult that state and validate the object before deciding to repeat work. Creating report-42 changes job identity and can produce another publication. If output is incomplete, recovery must account for effects already produced. Application design can provide ID-based lookup and safe repetition, but those properties need their own implementation and evidence. Do not automatically attribute them to the platform. Record what the client observed and what the service actually produced as distinct facts in the ticket. Resolve their difference through evidence rather than assuming cancellation.
Decide expansion and recovery at the right scope
In the fictional trial, canary has 12 failures in 100 requests and stable has 20 in 2000. The total 32/2100 is about 1.52%, but the approved rule stops expansion when canary exceeds 5% after at least 100 requests. Canary’s 12% triggers that decision. A 2% global alert serves another scope and does not replace the release criterion. After containing expansion, evaluate planned recovery and retain evidence for the affected revision. In Cloud Deploy, requesting rollback creates a new rollout based on an earlier release. Resource creation does not yet establish deployment success or functional recovery. Follow the outcome and validate the service. Restoring an application version also does not establish reversal of data changes or external effects; those elements need coverage in the recovery plan. Distinguish the request to recover from the evidence that recovery has actually completed.
Connect error budgets to verifiable actions
With a 99.5% SLO, permitted error rate is 0.5%. A window with 2% errors has a burn rate of 4: consumption proceeds at four times the tolerance. This does not mean 4% of the entire budget was spent. Remaining budget requires history, population and the SLO period. In the final exercise, calculate rates first, then write the action each signal justifies. If an alert reached a channel without an on-call operator, “improve alerts” is insufficient as a postmortem action. Define an owner, deadline and delivery rehearsal acknowledged by the intended operator, retaining the result. Closure should demonstrate correction of the failed path. These exercises use fictional data and documentation; they are neither cloud-platform measurements nor a bank’s internal procedures. The learning task is to connect every decision to what its evidence can establish.
Canary fails 12% of requests, but service totals show 1.52%; the decision depends on the revision’s approved criterion.
Common pitfalls
Averaging percentages without denominators, treating missing traces as nonexistent requests and treating timeout or rollout creation as final outcomes.
Related topics: SLOs and error budgets · Observability and diagnosis · Release acceptance
A signal needs scope, coverage and meaning; recovery is established only when the required outcome has been observed.
Reference: Alerting on SLOs · Current linked standard guide; edition date unconfirmed (2026-09-30 inspection)