← Professional Cloud DevOps Engineer: delivery and reliability
19 / 25 · 135 MIN

Delivery signals: compare versions and support promotion

Connect revision and telemetry, compare populations, interpret distributions and distinguish technical success from sufficient evidence.

1. Bind observations to the running version

During rollout of a fictional funds application, three teams may call delivery complete for different reasons. Platform confirms deployment, application owners confirm the process started, and support finds a dashboard without alerts. A decision to expand traffic needs to connect those facts to the revision serving requests and the outcomes produced. Before the window, define fields connecting release, environment, service, version and observations. Retain the query reference and interval so another person can repeat the evaluation. OpenTelemetry provides service.version to describe component version. Keep service.name as the logical name and check that the backend retains the version required for comparison. If the new binary still emits the old version, an absent candidate series does not prove absent execution. Correct the source and confirm new observations. Avoid relabelling all history because that also changes interpretation of genuinely old requests. deployment.environment.name describes environment but does not automatically select only production in a query. Define explicit filters and inspect observed values. A convincing legend does not fix a poorly selected population. Include a sample observation from each intended group in the evidence record, allowing the receiving support team to check that the mapping actually exists.

2. Compare populations with the same meaning

If candidate has ten errors in two hundred requests, its rate is 5%. Dividing those ten by two thousand requests across the service gives 0.5%, but changes the question. The numerator belongs to candidate while the denominator includes control. Likewise, summing series and dropping revision prevents versions from later being separated using that result. Decide which dimensions are needed before creating rules that retain only aggregates. Even with correct denominators, composition can mislead. In the exercise, control has 90 errors in 900 reads and one error in 100 writes. Candidate has 20 in 100 reads and 18 in 900 writes. Overall, control has 9.1% and candidate 3.8%. Within reads, however, rate rises from 10% to 20%; within writes, from 1% to 2%. Candidate receives proportionally more work from the less failure-prone operation. The aggregate improves while both operations worsen. Present the per-operation table and explain this difference to the promotion owner. These numbers are invented to teach comparison; they do not represent a real service or establish statistical significance. A useful review asks whether each comparison retains the workload meaning promised by the acceptance rule.

3. Bound time and confirm freshness

A candidate installed at 10:20 should not be assessed as though it produced every event since 09:30. Check exposure interval and the effective interval of each calculation. A dashboard can refresh now while still displaying a window containing earlier work. When data arrives late, separate event time from receipt time. In Cloud Logging, timestamp describes when an event occurred and receiveTimestamp when it was received. An event at 10:02 received at 10:12 does not become later than a 10:10 release, assuming the source recorded time correctly. In the laboratory, the rule requires an observation within the previous 120 seconds. At second 600, an observation from 450 is stale even if its report was re-exported at 590. This teaching limit should remain an explicit input. Define what happens when volume or coverage is missing too. Two error-free requests do not satisfy a rule requiring one hundred per operation and version. That does not establish functional failure; it establishes that required decision evidence is still absent. Keep this state distinguishable from approval and observed regression. The next team needs to know which data to collect to resolve uncertainty, rather than merely inherit an unexplained amber status.

4. Read distributions without inventing percentiles

Classic histogram buckets are cumulative. If 60 observations are at most 0.1 seconds, 90 at most 0.3 and 100 exist overall, 90% meet the 0.3 boundary. Adding the first two buckets counts faster observations again. Subtracting 60 from 90 answers a different question: how many are above 0.1 and at most 0.3. State units, boundaries and population before interpreting the result. A convention mismatch can turn an apparent regression into a calculation error. When combining instances, check measurement compatibility. If only A publishes boundary 0.3 while B publishes other boundaries, a numerator selected by le=0.3 omits B. Dividing by both instances’ total does not create the missing measurement. Nor does averaging each replica’s exported p95 yield global p95. The joint distribution requires aggregatable data. In PromQL, classic-histogram aggregation for histogram_quantile retains le and dimensions to keep separate, such as revision. Expressions in this block were reviewed against documentation; they were not executed against Prometheus. In real work, confirm types, labels, resets and compatibility before using a query as a promotion criterion. Keep a small known dataset available for checking whether the query answers the intended question.

5. Interpret an analysis job result

Cloud Deploy supports analysis jobs linked to Google Cloud Observability alert policies. When verify exists, analysis runs after it and before postdeploy. Configuration identifies policies and can restrict considered alerts using labels. Use the qualified policy identifier without appending a condition segment. If one policy covers multiple applications, confirm that the check selects its intended application. Do not reduce another service’s alerts to compensate for incorrectly scoped analysis. If duration ends without detected alerts, the job can succeed and rollout can continue. This behavior makes evidence design important. In a fictional case, the filter selects a revision that never received traffic. Technical success does not satisfy a local rule requiring candidate observations. Request a record of observed populations and freshness before declaring the version validated. Do not retrospectively change the criterion to justify the status color. If an exception is necessary, retain its reason, owner and accepted risk. A project manager can coordinate the decision but needs the technical team to confirm what signals actually measured. A correctly configured check still depends on appropriate instrumentation and data. Record those dependencies during preparation so that a missing observation has a defined investigation path during the window.

6. Use logs and traces to test hypotheses

Start with a hypothesis that evidence can contradict. If a query uses severity, inspect that field in the entry. A jsonPayload.level containing ERROR does not automatically populate severity in a direct API call. When analyzing exported logs, also confirm how the destination handles duplicates: insertId does not guarantee export deduplication. Counting repeated text as one event can remove distinct legitimate failures too. Document the chosen identity and preserve the ability to inspect examples. In traces, read structure and intervals before adding durations. A 100 ms parent can contain two overlapping 70 ms children. Their 140 ms sum is not parent wall-clock time. If a remote span appears to begin before the client sends, inspect synchronization and instrumentation before interpreting cross-clock differences as latency. Use parent IDs to guide investigation. For per-request correlation, avoid creating one metric series per UUID. Unbounded labels multiply series; hashing identifiers does not bound the number of values. Retain controlled dimensions in metrics and use appropriately accessible logs or traces to locate specific requests. Explain which observations support a hypothesis and which remain ambiguous, so another engineer can challenge the conclusion without losing the investigation’s context.

7. Protect decisions from incomplete conclusions

A relative comparison needs operational context. If candidate and control share a database, candidate can degrade both. A small between-group difference does not mean the service is healthy. Keep appropriate absolute criteria alongside comparison. In the example, both rise from 1% to 8% errors. A rule accepting differences below one percentage point misses the common degradation. Investigate the dependency and manage impact before expanding exposure. Also check whether indicator definition stayed constant. If new instrumentation removes timeouts from the numerator while retaining them in the total, rate falls without establishing improvement. Use the agreed failure definition and explain changes before comparing. At the review meeting, present population, interval, criterion, result and known gaps. Avoid reducing everything to a color. A useful decision identifies the next action: collect missing data, fix instrumentation, hold promotion or present sufficient evidence to the owner. Retain the query and counts used. These practices help APS support a decision under deadline pressure and allow the project to explain why it advanced or remained on hold. They also preserve the distinction between an observed regression and a claim about its cause, which may require further investigation.

8. Per-operation evidence laboratory

The Python code uses a fictional rule: reads and writes need complete, fresh observations with at least one hundred requests per version. Within each operation, candidate rate must not exceed control rate. Freshness limit is 120 seconds, including the boundary. The program validates nonnegative integer counts and rejects errors exceeding requests. It uses cross multiplication to compare rates and Fraction to display exact values. This is neither a statistical significance test nor a Cloud Deploy implementation. Run the eight cases and start with mixed-population. Its aggregate favors candidate, but both operations regress and the decision is hold. too-few-writes, stale, incomplete and missing-operation remain inconclusive. freshness-boundary shows that an age exactly equal to the limit is accepted by the rule. zero-requests does not become a zero error rate. The program checks 625 rate combinations, sixteen quality combinations and ten invalid inputs. promotionAuthorized stays false in every case: eligible-for-review describes only evidence meeting this rule. Then change one count and predict the decision before running. Explain which check changed and what investigation a real service would need. The exercise makes no cloud queries, executes no PromQL and authorizes no deployments. Its output is a worksheet for discussing evidence, not a replacement for an approved operational process.

"""Original fictional review rule, not Cloud Deploy or a statistical canary test."""
from copy import deepcopy
from fractions import Fraction
import hashlib
import itertools
import json
from pathlib import Path


def integer(value, label):
 if type(value) is not int or value < 0:
 raise ValueError(label + ' must be a nonnegative integer')
 return value


def ratio(errors, requests):
 return str(Fraction(errors, requests)) if requests else None


def evaluate(data, now=600, max_age=120, minimum=100):
 for key, value in [('now', now), ('max_age', max_age), ('minimum', minimum)]:
 integer(value, key)
 if minimum == 0 or not isinstance(data, dict) or set(data)!= {'observedAt', 'complete', 'groups'}:
 raise ValueError('invalid rule or evidence shape')
 observed = integer(data['observedAt'], 'observedAt')
 if observed > now or type(data['complete']) is not bool or not isinstance(data['groups'], dict):
 raise ValueError('invalid evidence metadata')
 expected = {'read', 'write'}
 if not set(data['groups']).issubset(expected):
 raise ValueError('unknown operation')
 reasons, comparisons = [], []
 if now - observed > max_age:
 reasons.append('stale')
 if not data['complete']:
 reasons.append('incomplete')
 if set(data['groups'])!= expected:
 reasons.append('missing-operation')
 totals = {name: [0, 0] for name in ['control', 'candidate']}
 for operation, group in sorted(data['groups'].items):
 if not isinstance(group, dict) or set(group)!= {'control', 'candidate'}:
 raise ValueError('each operation needs control and candidate')
 for name, values in group.items:
 if not isinstance(values, dict) or set(values)!= {'requests', 'errors'}:
 raise ValueError('invalid counts shape')
 n = integer(values['requests'], 'requests')
 errors = integer(values['errors'], 'errors')
 if errors > n:
 raise ValueError('errors exceed requests')
 if n < minimum:
 reasons.append('low-volume:' + operation + ':' + name)
 totals[name][0] += errors
 totals[name][1] += n
 a, b = group['control'], group['candidate']
 regression = None if not a['requests'] or not b['requests'] else b['errors'] * a['requests'] > a['errors'] * b['requests']
 comparisons.append({'operation': operation, 'controlRate': ratio(a['errors'], a['requests']),
 'candidateRate': ratio(b['errors'], b['requests']), 'regression': regression})
 decision = 'inconclusive' if reasons else ('hold' if any(x['regression'] for x in comparisons) else 'eligible-for-review')
 return {'decision': decision, 'reasons': reasons, 'operations': comparisons,
 'aggregateRates': {k: ratio(*v) for k, v in totals.items},
 'promotionAuthorized': False}


def evidence(read=(1, 100, 1, 100), write=(1, 100, 1, 100)):
 result = {'observedAt': 600, 'complete': True, 'groups': {}}
 for name, (ce, cn, ne, nn) in [('read', read), ('write', write)]:
 result['groups'][name] = {'control': {'errors': ce, 'requests': cn}, 'candidate': {'errors': ne, 'requests': nn}}
 return result


def main:
 base = evidence
 fixtures = [('equal', deepcopy(base), 'eligible-for-review'),
 ('mixed-population', evidence((90, 900, 20, 100), (1, 100, 18, 900)), 'hold')]
 low = evidence(write=(1, 100, 0, 2)); fixtures.append(('too-few-writes', low, 'inconclusive'))
 stale = deepcopy(base); stale['observedAt'] = 450; fixtures.append(('stale', stale, 'inconclusive'))
 partial = deepcopy(base); partial['complete'] = False; fixtures.append(('incomplete', partial, 'inconclusive'))
 missing = deepcopy(base); del missing['groups']['write']; fixtures.append(('missing-operation', missing, 'inconclusive'))
 boundary = deepcopy(base); boundary['observedAt'] = 480; fixtures.append(('freshness-boundary', boundary, 'eligible-for-review'))
 empty = evidence(write=(0, 100, 0, 0)); fixtures.append(('zero-requests', empty, 'inconclusive'))
 output = []
 for name, data, decision in fixtures:
 before = deepcopy(data)
 result = evaluate(data)
 assert data == before and result['decision'] == decision
 output.append({'id': name, **result})
 assert output[1]['aggregateRates'] == {'control': '91/1000', 'candidate': '19/500'}
 assert all(x['regression'] for x in output[1]['operations'])
 comparisons = 0
 for cr, nr, cw, nw in itertools.product(range(5), repeat=4):
 data = evidence((cr, 100, nr, 100), (cw, 100, nw, 100))
 result = evaluate(data)
 assert result['decision'] == ('hold' if nr > cr or nw > cw else 'eligible-for-review')
 comparisons += 1
 quality_cases = 0
 for fresh, complete, enough, present in itertools.product([False, True], repeat=4):
 data = evidence(write=(1, 100, 0, 100 if enough else 2))
 data['observedAt'] = 600 if fresh else 450
 data['complete'] = complete
 if not present:
 del data['groups']['write']
 result = evaluate(data)
 assert result['decision'] == ('eligible-for-review' if fresh and complete and enough and present else 'inconclusive')
 quality_cases += 1
 invalid = []
 for field, value in [('observedAt', 601), ('observedAt', True), ('observedAt', -1), ('complete', 1)]:
 data = deepcopy(base); data[field] = value; invalid.append(data)
 for field, value in [('requests', -1), ('requests', True), ('requests', 1.5), ('errors', 101)]:
 data = deepcopy(base); data['groups']['read']['candidate'][field] = value; invalid.append(data)
 bad = deepcopy(base); bad['groups']['extra'] = {}; invalid.append(bad)
 bad = deepcopy(base); del bad['groups']['read']['control']; invalid.append(bad)
 for data in invalid:
 try:
 evaluate(data)
 except ValueError:
 pass
 else:
 raise AssertionError('invalid evidence accepted')
 print(json.dumps({'scriptSha256': hashlib.sha256(Path(__file__).read_bytes).hexdigest,
 'fixtures': output, 'rateComparisons': comparisons, 'qualityCombinations': quality_cases,
 'invalidInputs': len(invalid), 'inputPreserved': True, 'cloudExecuted': False,
 'promqlExecuted': False, 'network': False, 'persistentWrites': False,
 'limitations': 'Fictional deterministic counts rule; no statistical significance, vendor execution, deployment authority or causal proof.'}, indent=2))


if __name__ == '__main__':
 main
IN PRACTICE

Candidate has 3.8% overall errors versus control’s 9.1%, but its read and write error rates doubled.

Common pitfalls

Mixing revisions, treating export time as observation, averaging percentiles or accepting no alerts without traffic.

Related topics: Instrumentation and telemetry transport · Continuity and outcome reconciliation · Performance and cost per completed work

Take this idea with you

A delivery decision needs attributable, comparable and fresh data; technical results must be interpreted within the agreed criterion.

Create account

Reference: Professional Cloud DevOps Engineer exam guide · Current linked guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.