← Professional Cloud DevOps Engineer: delivery and reliability
24 / 25 · 135 MIN

Diagnosis with incomplete telemetry

Bound what logs, metrics, traces and probes establish, identify coverage gaps and test hypotheses during production incidents.

1. Separate service outcome from observation quality

A fictional APS team receives three signals: one region loses metrics, probes fail and the dashboard shows zero errors. Before concluding recovery or total failure, identify what each observation actually measures. Zero errors among observed requests differs from zero observed requests. A region without data might be unavailable, isolated from monitoring or outside query scope. These hypotheses lead to different actions. Build a short record containing user journey, expected population, interval, source, monitor dependencies, observation and inference limit. For example: a positions query works through the alternative path at 10:12; reconciliation has no recent result. That record preserves partial success without inventing the batch outcome. The example does not represent internal BNP Paribas procedures. Choose the next test for its ability to distinguish hypotheses. If probes share a proxy, an authorized path with different dependencies can help locate the problem. Define a test's scope and expected result before execution. During an incident, changing application, DNS and monitoring simultaneously can destroy the comparison needed to understand the cause. Record the sequence of changes and keep what remains unobserved explicit.

2. Confirm scope before interpreting absence

The central monitoring project shows only projects included in the applicable scope. If series exist in funds-prod but that project is absent from the metrics scope being used, the incomplete central view does not establish collection failure. Confirm project, resources, labels and interval before reinstalling agents. Comparing a source query with the central query helps locate the difference. For logs, a view controls the visible subset and IAM governs access. Two operators can obtain different results because they use different views or permissions. Request necessary access to the relevant scope through the authorized process. Do not expand permissions indiscriminately or turn an empty query into a claim that no events occurred. Then follow the routing path. A sink needs appropriate filters and its writer identity needs destination permission. Successful publishing by a personal account does not validate the sink identity. Exclusion in one independent sink does not imply exclusion in every other sink. A new sink does not automatically route entries received before it was created. If history is needed, identify where it remains stored and the appropriate copying mechanism. Fixing current export does not establish that the earlier interval was recovered.

3. Distinguish zero,absence and partial coverage in PromQL

In Prometheus, up=0 means the scrape failed. It does not directly measure business requests or identify the cause of failure. A target can respond to users while being inaccessible to the scraper. Another situation exists: a target leaves discovery and its series become stale. When they disappear from an instant query, a panel counting only zeros can omit that instance entirely. Keep expected inventory separate from the set that responded. In this lesson's original situation, 100 shards are expected and only 80 report success. You can state success for those observed and coverage of 80% of inventory. You cannot publish 100% success for all of them or call the remaining 20 confirmed failures. Their state remains undetermined. The absent_over_time function indicates absence of the selection over the window. A broad selection containing A samples does not automatically enumerate missing B. Similarly, adding or vector(0) to an empty expression can produce a zero without actual observation. Before using dashboard fill values, define what zero means and keep collection quality visible. Expressions in this lesson are teaching examples; the local exercise does not run a Prometheus server or replace query validation in the actual environment.

4. Follow data through the Collector

Effective OpenTelemetry Collector configuration must connect components through pipelines in the service section. Declaring a receiver in a file does not establish that it is active. Check that the component is referenced in the correct pipeline and that running configuration matches the reviewed configuration. Do not rely solely on a filename or dashboard existence to confirm collection. Then observe flow boundaries: items accepted and refused by receivers, processor input and output, items sent or rejected by exporters and queue occupancy. Metric names and availability depend on version and internal-telemetry configuration; confirm them for the installed binary. Compare the same signal type and interval. Received spans and exported batches are not directly additive units. In the controlled example, 1,000 spans enter and a deliberate filter removes 200; exporting 800 is consistent. In another case, the receiver refuses items while the exporter sends everything it receives. Exporter success leaves the fate of refused items undetermined. Investigate source retry behavior and pipeline limits. A reachable health extension does not by itself establish telemetry delivery to the destination. Use bounded synthetic data to follow the path without exposing actual customer payloads.

5. Explain what probes actually test

A probe has a protocol, destination, network path, identity and success criterion. In a public HTTP uptime check, redirects are followed and criteria apply to the final response. If /portal redirects to /maintenance and only 200 is required, the check can pass without demonstrating portal functionality. An uptime check also does not automatically execute page JavaScript. Choose a functional test appropriate to the requirement you intend to observe. TCP connectivity to port 443 does not prove TLS validation. That requirement needs a check exercising the protocol and certificate criteria. The consulted documentation states that private uptime checks disable SSL validation regardless of configuration. Do not use their result as proof of certificate validity. Record that limitation in the coverage design. Also consider monitor identity and dependencies. An expired synthetic account can fail while valid accounts remain functional. Three locations using the same proxy are not three observations proven independent of that dependency. An alternative path helps investigation but does not erase users who still depend on the proxy. The report should identify which journeys were observed and where evidence remains incomplete.

6. Correlate signals without inventing completeness

A retained trace shows only what was instrumented, propagated, collected and retained. A missing client span does not prove an external call never occurred. Confirm instrumentation, sampling and export before interpreting the gap. Seek complementary evidence with consistent identifiers and intervals. An application log recording an attempt and a trace lacking the span can suggest incomplete coverage without yet establishing the technical cause. VPC Flow Logs is not a capture of every packet either. Secondary sampling at 100% retains flow logs produced by primary sampling; it does not remove earlier sampling. An empty search does not establish that a communication attempt never existed. Check interfaces and scope, interval, filters and collection limits before drawing network conclusions. In a fictional incident, an operator suspects DNS but has only a timeout and no flow log. Form hypotheses that can be distinguished: resolution, connection, TLS negotiation and functional response. Choose observations that locate the failure boundary. Keep the test bounded to the question being investigated and record useful negative results. A plausible hypothesis should remain a hypothesis until supporting evidence exists; multiplying dashboards backed by the same source does not add independence.

7. Validate detection,notification and handover

A useful alert requires more than a created channel. The condition must select the right resources, receive appropriate data, evaluate intended behavior and produce a notification reaching the responsible person. A webhook connection test proves part of transport. It does not prove the actual policy is linked to the channel or its threshold can trigger with the selected data. Prepare an authorized rehearsal with a controlled condition, agreed destination and clear test identification. Observe the complete path to the receiver and confirm who owns the action. Do not run that rehearsal against actual customers without the appropriate change process. Documentation includes test options specific to some channels; do not assume a universal button validates every policy. A snooze window suppresses alerts and notifications in its defined scope. Silence during that window does not establish normal service operation. Record scope, duration, owner and planned alternative observation. If maintenance ends early, review suppression timing. At handover, separate observed failures, missing coverage, hypotheses under investigation and pending actions. A report identifying the reconciliation gap is more useful than a global green status that nobody can connect to a recent functional test.

8. Exercise: coverage by journey and declared group

The Python program receives expected journey inventory and a snapshot containing each probe's latest observation. Every record has a unique ID, journey, declared dependency group, observation time and pass, fail or unknown result. The clock is fictional and shared. The teaching rule requires two distinct groups with known outcomes per journey and age between zero and 120 seconds, inclusive. A group is an inventory declaration, not proof of physical independence. Five probes using proxy-a count as one group. None for time or group preserves uncertainty; an unknown result does not provide known coverage either. Stale or future records are excluded from fresh coverage. A recent failure is retained even when its group is unknown. The program does not turn gaps into success or hide failures when another journey lacks data. Run the examples and modify the snapshot: remove reconciliation, repeat groups, use observedAt=480 with now=600 and compare with 479 and 620. Explain results before looking at output. Then add a probe contradicting another in the same group. The algorithm retains both observations and counts the group once. CoverageComplete means only satisfying the declared rule. The program makes no requests, executes no PromQL and proves neither service health nor production authorization.

"""Original offline snapshot exercise, not a monitor or availability estimator.

One latest observation per probe ID. Groups are declared dependency boundaries,
not discovered or proven independence. Ages use one fictional shared clock.
"""
import copy
import hashlib
import itertools
import json
from pathlib import Path

def name(value):
 if type(value) is not str or not value.strip:
 raise ValueError('nonempty label required')
 return value

def integer(value, minimum=0):
 if type(value) is not int or value < minimum:
 raise ValueError('integer outside permitted range')
 return value

def evaluate(journeys, observations, now, max_age, minimum_groups):
 integer(now); integer(max_age); integer(minimum_groups, 1)
 if type(journeys) is not list or not journeys:
 raise ValueError('nonempty journey inventory required')
 for j in journeys: name(j)
 if len(journeys)!= len(set(journeys)):
 raise ValueError('duplicate journey')
 if type(observations) is not list:
 raise ValueError('snapshot list required')
 ids = set
 validated = []
 for row in observations:
 if type(row) is not dict or set(row)!= {'id','journey','dependencyGroup','observedAt','result'}:
 raise ValueError('observation fields')
 name(row['id']); name(row['journey'])
 if row['id'] in ids: raise ValueError('duplicate probe ID')
 ids.add(row['id'])
 if row['journey'] not in journeys: raise ValueError('unexpected journey')
 if row['dependencyGroup'] is not None: name(row['dependencyGroup'])
 if row['observedAt'] is not None: integer(row['observedAt'])
 if row['result'] not in ('pass','fail','unknown'): raise ValueError('observation result')
 validated.append(row)
 output = []
 for journey in sorted(journeys):
 relevant = sorted((r for r in validated if r['journey'] == journey),key=lambda r:r['id'])
 groups, issues, successes, failures = {}, [], [], []
 for row in relevant:
 reason = None
 if row['observedAt'] is None: reason = 'unknown-time'
 elif row['observedAt'] > now: reason = 'future-time'
 elif now-row['observedAt'] > max_age: reason = 'stale'
 if reason:
 issues.append({'id':row['id'],'reason':reason})
 continue
 if row['result'] == 'unknown':
 issues.append({'id':row['id'],'reason':'unknown-result'})
 continue
 (successes if row['result'] == 'pass' else failures).append(row['id'])
 group = row['dependencyGroup']
 if group is None:
 issues.append({'id':row['id'],'reason':'unknown-group'})
 else:
 groups.setdefault(group,set).add(row['result'])
 output.append({'journey':journey,'freshGroups':[
 {'group':g,'results':sorted(groups[g])} for g in sorted(groups)],
 'requiredGroups':minimum_groups,'coverageMet':len(groups)>=minimum_groups,
 'observedPasses':successes,'observedFailures':failures,'observationIssues':issues})
 return {'journeys':output,'coverageComplete':all(j['coverageMet'] for j in output),
 'failureObserved':any(j['observedFailures'] for j in output),
 'hasObservationIssues':any(j['observationIssues'] for j in output),
 'independenceProven':False,'serviceHealthProven':False,'productionAuthorized':False}

def baseline:
 return ['query','reconciliation'], [
 {'id':j+'-'+g,'journey':j,'dependencyGroup':g,'observedAt':590,'result':'pass'}
 for j in ('query','reconciliation') for g in ('a','b')]

def exercise:
 fixtures=[]
 def record(label, mutate):
 j,o=baseline;mutate(j,o);before=copy.deepcopy((j,o));r=evaluate(j,o,600,120,2)
 assert (j,o)==before
 fixtures.append({'id':label,**r});return r
 assert record('complete-pass',lambda j,o:None)['coverageComplete']
 assert not record('missing-journey',lambda j,o:o.__delitem__(slice(2,None)))['coverageComplete']
 assert not record('shared-group',lambda j,o:[r.update(dependencyGroup='a') for r in o])['coverageComplete']
 assert record('boundary-fresh',lambda j,o:o[0].update(observedAt=480))['coverageComplete']
 assert not record('just-stale',lambda j,o:o[0].update(observedAt=479))['coverageComplete']
 assert not record('future-time',lambda j,o:o[0].update(observedAt=620))['coverageComplete']
 assert not record('unknown-time',lambda j,o:o[0].update(observedAt=None))['coverageComplete']
 assert not record('unknown-group',lambda j,o:o[0].update(dependencyGroup=None))['coverageComplete']
 assert not record('unknown-result',lambda j,o:o[0].update(result='unknown'))['coverageComplete']
 f=record('failure-and-gap',lambda j,o:(o[0].update(result='fail'),o.__delitem__(slice(2,None))))
 assert f['failureObserved'] and not f['coverageComplete']
 f=record('mixed-group',lambda j,o:o.append({'id':'query-extra','journey':'query','dependencyGroup':'a','observedAt':590,'result':'fail'}))
 assert f['failureObserved'] and f['coverageComplete']
 f=record('unknown-group-failure',lambda j,o:o[0].update(result='fail',dependencyGroup=None))
 assert f['failureObserved'] and not f['coverageComplete']
 combinations=0
 for states in itertools.product(('pass','fail','unknown','stale','future'),repeat=4):
 j,o=baseline
 for row,state in zip(o,states):
 row['result']=state if state in ('pass','fail','unknown') else 'pass'
 row['observedAt']=479 if state=='stale' else 620 if state=='future' else 590
 r=evaluate(j,o,600,120,2)
 assert r['coverageComplete']==all(x in ('pass','fail') for x in states)
 assert r['failureObserved']==('fail' in states)
 assert r['hasObservationIssues']==any(x not in ('pass','fail') for x in states)
 combinations+=1
 j,o=baseline; expected=evaluate(j,o,600,120,2); permutations=0
 for order in itertools.permutations(o):
 assert evaluate(list(reversed(j)),list(order),600,120,2)==expected
 permutations+=1
 invalid=[
 lambda j,o:j.clear, lambda j,o:j.append('query'),lambda j,o:j.append(''),
 lambda j,o:o.append(copy.deepcopy(o[0])),lambda j,o:o[0].update(extra=1),
 lambda j,o:o[0].pop('result'),lambda j,o:o[0].update(journey='payments'),
 lambda j,o:o[0].update(id=' '),lambda j,o:o[0].update(dependencyGroup=''),
 lambda j,o:o[0].update(observedAt=-1),lambda j,o:o[0].update(observedAt=True),
 lambda j,o:o[0].update(result='healthy'),lambda j,o:o[0].update(observedAt=590.5),
 ]
 for mutate in invalid:
 j,o=baseline;mutate(j,o)
 try:evaluate(j,o,600,120,2)
 except ValueError:pass
 else:raise AssertionError('invalid snapshot accepted')
 settings=[(True,120,2),(600,-1,2),(600,120,0),(600,120,False)]
 for values in settings:
 j,o=baseline
 try:evaluate(j,o,*values)
 except ValueError:pass
 else:raise AssertionError('invalid settings accepted')
 return {'scriptSha256':hashlib.sha256(Path(__file__).read_bytes).hexdigest,
 'fixtures':fixtures,'snapshotCombinations':combinations,'orderPermutations':permutations,
 'invalidInputs':len(invalid)+len(settings),'inputPreserved':True,
 'network':False,'persistentWrites':False,'cloudExecuted':False,'promqlExecuted':False}

if __name__=='__main__':
 print(json.dumps(exercise,ensure_ascii=False,sort_keys=True,indent=2))
IN PRACTICE

Two groups observe query, but none observes reconciliation. Query success is recorded and reconciliation retains an explicit gap without being classified as success or failure.

Common pitfalls

Replacing absence with zero; counting probes as independent groups; equating scrape success with functional success; treating notification silence as recovery.

Related topics: Synthetic monitoring and SLIs · Telemetry pipelines and access control · Incident diagnosis and communication

Take this idea with you

Connect every conclusion to observation scope, age and limits. Preserve known failures and gaps together until sufficient evidence exists.

Create account

Reference: Professional Cloud DevOps Engineer exam guide · Current linked guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.