← Professional Cloud DevOps Engineer: delivery and reliability
14 / 15 · 135 MIN

Observability: populations and diagnosis

Investigate telemetry gaps, interpret queries and alerts, and use traces with explicit context and population.

1. Follow the evidence path

A fictional team monitors a positions-query service. The application still responds, but traces have stopped arriving. Draw the path before changing configuration: application, transport, receiver, processing, exporter and destination. At each boundary, find an observation that distinguishes input, transformation and delivery. A running Collector process does not prove the application can reach it. A growing receive counter does not prove the destination accepted data. An empty dashboard can result from an incorrect query even while delivery works. In the first trial, the application logs OTLP connection failures and reception does not increase. Investigate the receiver endpoint, protocol and accessibility. In the second, reception increases but the exporter logs PermissionDenied. Check effective identity and authorization at the configured destination. Increasing replicas or sampling does not fix that rejection. To validate the correction, generate an identifiable synthetic sample, follow it through the path and confirm arrival. Record the trial time and identifier. The APS owner should be able to explain which segment has been demonstrated and which still requires investigation, avoiding presenting one component’s health as proof of the complete chain.

2. Size queues and contain diagnostics

An empty queue with room for 240 batches receives 12 batches/s and cannot export. In the simplified model, it fills in 20 seconds. This calculation uses batches, not spans or bytes, and assumes no other queues, departures or expiry during the interval. Before applying the reasoning to a real Collector, confirm the queue unit and configuration for the exporter in use. Monitor occupancy, refusals, retries and available storage. A larger queue buys time; recovery still requires sustainable output and backlog handling. Persistence on durable storage can protect queued data across restarts, but creates neither infinite space nor recovery of data lost before writing. In the change plan, state what may happen to pending data and how delivery will be checked after restart. If payload inspection is necessary, use synthetic data with the same structure and a controlled diagnostic output. Avoid copying account numbers into broadly accessible logs. Base64 is reversible and does not resolve that exposure. This lesson’s local exercise calculates over a fictional inventory; it does not run a Collector, implement persistence or validate delivery of real telemetry.

3. Define probe and audit coverage

An HTTP 200 response from /health can coexist with a failed authenticated query. Define the functional journey that matters to the business: authenticate a test identity, query synthetic data and confirm a known result. Establish cleanup steps when a trial creates state. Merely increasing /health frequency improves observation of that endpoint without adding functional coverage. At RUN handover, document what is tested, from where, under which identity and how to distinguish application unavailability from an expired test credential. For auditing, first confirm that the event is generated. A DATA_READ operation in a service that does not log it by default needs the applicable configuration enabled. Admin Activity does not replace every data read. Also inspect exemptions, inherited configuration, routing and reading permissions before concluding that nobody accessed data. Validate a new authorized operation and its corresponding record. A sink does not retroactively create audit events that were never generated. These examples establish neither legal obligations nor any bank’s policy: they help formulate concrete technical questions for security and operations owners.

4. Query and preserve without inventing data

Create three synthetic entries: result="ok", result="failed" and one without result. In Cloud Logging, jsonPayload.result!="ok" does not include the missing field. NOT (jsonPayload.result="ok") includes it because it negates a comparison that fails in that case. Write the expected result before running a query. Also check parentheses: in this language OR takes precedence over AND. If the intention is (A AND B) OR C, retain that explicit expression. Do not automatically carry SQL rules into another query engine. When filtering logName by equality, use the observed identifier, including %2F when a slash belongs to the encoded log name. No results also require checking time window, scope and preservation. Increasing retention from 30 to 90 days does not recreate permanently deleted entries without a copy. Record the gap and decide how future evidence will be preserved. Locking a log bucket is irreversible; a plan promising to unlock it the following week needs correction before execution. Separate that decision from sink creation and reading permissions. The lesson examples are interpretation exercises; the Python lab executes neither Cloud Logging queries nor retention changes.

5. Read the complete alert rule

A policy can combine high CPU on one VM with high memory on another. If the requirement is to find both on the same VM, use resource-matching semantics and preserve compatible labels. Plain AND and AND_WITH_MATCHING_RESOURCE do not represent the same condition. Also check the retest window: four minutes above threshold, one aligned measurement below, then three minutes above do not satisfy six continuous minutes. The healthy sample resets the count. The policy should respond to the intended operational condition, not merely produce a notification during a convenient trial. Closure also needs interpretation. When missing data is treated as non-violating, a collection failure can close an alert without proving application recovery. Monitor collection health and document that choice. For cumulative metrics in PromQL, calculate rate per series before summing, preserving detection of individual resets. Confirm the result’s unit in the dashboard. When investigating low global CPU alongside a saturated zone, recover zone grouping and relate it to routing and load. Global reduction removes distinctions that may be essential to the decision.

6. Exercise: different windows and populations

Save the code as run.py and run python3 run.py with Python 3.13. The model uses Fraction to compare exact values. A 99.8% target allows 0.2% errors; observing 1% produces a burn rate of five. The teaching rule requires values greater than six in both windows. Predict both_high, short_recovered and equal_threshold before reading the JSON. If an input is unknown, this model returns null: it does not reproduce a real service’s missing-data policy. Compare the 36 window combinations checked by the program. The second trial has 200 errors among 20,000 requests but retains every error and only 300 successful traces. The complete rate is 1%; the selected sample has 40%. The program varies retained successes across 101 cases while keeping the original population fixed. Confirm that more detail per trace does not correct unequal selection. Finally, check retesting, queue capacity and p95: 9,900 requests taking 100 ms and 100 taking 20 seconds still have p95 of 100 ms under nearest-rank calculation. The lab checks arithmetic and explicit decisions without executing PromQL, Cloud Monitoring or a sampler.

7. Correlate traces without losing context

If A calls B and both export spans but B always creates an independent trace, check context injection and extraction on the call. Giving spans matching names does not establish a parent relationship. Using a fixed identifier for every request mixes executions. To link an entry sent directly through the Logging API to an existing trace, populate trace with the actual identifier. In the inspected documentation, simple TRACE_ID is preferred; projects/PROJECT_ID/traces/TRACE_ID remains documented as legacy. A business UUID in jsonPayload.request_id can support searching but does not automatically replace that link. Also define boundaries for baggage. Context can reach external services through instrumented clients; session tokens and account data should not be included in the proposed example. Use approved identifiers and check the actual propagation path. When analyzing traces retained for errors or latency, document sampling criteria before calculating overall percentages. Detailed evidence helps investigate particular requests; SLI estimation requires a suitable population. At handover, include trace origin, selection policy, correlation fields and owners for each service.

8. Investigate hypotheses and deliver useful conclusions

A release coincides with a routing change and latency rises. An automated interpretation attributes the failure to the release. Record that conclusion as a hypothesis and seek observations distinguishing version, region and dependency. Plan a controlled mitigation with a validation criterion. Changing everything simultaneously and attributing improvement to one change does not produce a sound explanation. Restoring service remains the incident priority, while preserving action sequence and separating observations from inference is still possible. Prepare the report around three concrete questions: which population was measured, which part of the path was demonstrated and what decision follows? In the in-memory queue case, describe risk before restart and verify new delivery afterward. In the selected-trace case, retain the complete counter’s rate and use the sample for diagnosis. Practise a business update: “The error counter covers all requests; retained traces are selected for diagnosis and use a different denominator.” The operational summary is to follow boundaries, confirm generation, query with correct semantics, interpret the alert rule and communicate coverage limits. These criteria connect observability with SRE, incident management and retention cost.

"""Original local observability arithmetic; no query engine or alert service."""
from fractions import Fraction
from hashlib import sha256
from itertools import product
from pathlib import Path
import json
import platform


def ratio(errors, total):
 if type(errors) is not int or type(total) is not int:
 raise ValueError('counts must be integers, not booleans')
 if not 0 <= errors <= total:
 raise ValueError('counts must satisfy 0 <= errors <= total')
 return Fraction(errors, total) if total else None


def burn(errors, total, target=Fraction(998, 1000)):
 if not isinstance(target, Fraction) or not 0 < target < 1:
 raise ValueError('target must be an exact fraction between zero and one')
 value = ratio(errors, total)
 return value/(1-target) if value is not None else None


def combined(long_rate, short_rate, threshold=Fraction(6)):
 # Unknown input leaves this educational decision unevaluable.
 # This is not a claim about a configured vendor missing-data policy.
 if long_rate is None or short_rate is None:
 return None
 return long_rate > threshold and short_rate > threshold


def consecutive_high(values):
 longest = current = 0
 for value in values:
 if type(value) is not bool:
 raise ValueError('aligned condition inputs must be booleans')
 current = current+1 if value else 0
 longest = max(longest, current)
 return {'longest': longest, 'current': current}


def run:
 observed = burn(100, 10000)
 assert observed == 5
 fixtures = {
 'both_high': combined(Fraction(8), Fraction(8)),
 'short_recovered': combined(Fraction(8), Fraction(2)),
 'long_recovered': combined(Fraction(2), Fraction(8)),
 'equal_threshold': combined(Fraction(6), Fraction(8)),
 'unknown_short': combined(Fraction(8), None),
 'zero_population': burn(0, 0),
 }
 assert fixtures == dict(both_high=True, short_recovered=False, long_recovered=False,
 equal_threshold=False, unknown_short=None, zero_population=None)
 combinations = 0
 for long_rate, short_rate in product([None, Fraction(0), Fraction(5), Fraction(6), Fraction(7), Fraction(10)], repeat=2):
 expected = None if long_rate is None or short_rate is None else min(long_rate, short_rate) > 6
 assert combined(long_rate, short_rate) == expected
 combinations += 1
 complete = ratio(200, 20000)
 selected = ratio(200, 500) # 200 errors + 300 selected successes, fictional inventory.
 assert complete == Fraction(1,100) and selected == Fraction(2,5)
 selection_checks = 0
 for selected_successes in range(0, 19801, 198):
 sample = ratio(200, 200+selected_successes)
 assert sample >= complete
 assert (sample == complete) == (selected_successes == 19800)
 selection_checks += 1
 retest = consecutive_high([True]*4+[False]+[True]*3)
 assert retest == {'longest':4, 'current':3}
 assert Fraction(240, 12) == 20
 ordered = [100]*9900+[20000]*100
 rank = (95*len(ordered)+99)//100
 assert rank == 9500 and ordered[rank-1] == 100 and max(ordered) == 20000
 invalid = 0
 for args in [(-1,10),(11,10),(1,-1),(True,10),(1,False),(1.5,10),(1,'10')]:
 try:
 ratio(*args)
 except ValueError:
 invalid += 1
 assert invalid == 7
 target_invalid = 0
 for target in [Fraction(0),Fraction(1),Fraction(-1),0.998]:
 try:
 burn(1,100,target)
 except ValueError:
 target_invalid += 1
 assert target_invalid == 4
 return {'scriptSha256':sha256(Path(__file__).read_bytes).hexdigest,
 'pythonVersion':platform.python_version,'fixtures':fixtures,
 'burnRate':str(observed),'populationErrorRate':str(complete),
 'selectedErrorRate':str(selected),'windowCombinations':combinations,
 'selectionChecks':selection_checks,'retest':retest,'queueSeconds':20,
 'nearestRankP95Milliseconds':100,'maximumMilliseconds':20000,
 'invalidCountCases':invalid,'invalidTargetCases':target_invalid,
 'network':False,'persistentWrites':False,'vendorExecution':False,
 'independentVerification':False,
 'limitations':'Fictional exact counts and precomputed window rates. No Cloud Monitoring, PromQL, collector, sampling implementation, late-data model or production alert lifecycle.'}


if __name__ == '__main__':
 print(json.dumps(run, ensure_ascii=False, indent=2))
IN PRACTICE

A complete counter shows 1% errors while selected traces show 40%; both observations can be correct.

Common pitfalls

Confusing reception with delivery, closure with recovery, a sample with its population and timing correlation with confirmed cause.

Related topics: SRE and error budgets · FinOps and telemetry preservation · Incident management and communication

Take this idea with you

Observability supports a decision when the path, population and interpretation rule are clear.

Create account

Reference: Troubleshooting the Collector · Current linked guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.