← Professional Cloud DevOps Engineer: delivery and reliability
22 / 25 · 135 MIN

Pipeline continuity and recovery compatibility

Prepare recovery options spanning artifacts, data, configuration, messages and inference, with explicit evidence for each contract.

1. Define the recovery unit

A recoverable release needs an identifiable combination of code, configuration, data and dependencies. In a fictional banking reconciliation exercise, version A worked with schema s1, secret 7 and m1 messages. Version B introduced s2 and m2. Retaining image A identifies code; the team must still establish that the current service satisfies A's contracts. This example does not describe internal BNP Paribas procedures. Build a record for each recovery option: image digest, versioned configuration, secret reference without its value, accepted schemas, emitted and consumed formats, required endpoints, tests and decision owner. Also record what can keep running during the change. A monthly batch and a lagging consumer can retain dependencies absent from the web dashboard. Use three questions in the preparation meeting: can we obtain the combination, can it interpret current state and can it serve expected demand? An unknown answer should produce an evidence collection action with an owner. Do not turn it into yes under window pressure. The result is a demonstrable option, with constraints and stop criteria, which the next shift can understand without reconstructing earlier conversations.

2. Retain and retrieve concrete artifacts

A tag such as stable can point today to different content from the rehearsal. Record the observed digest and full location. Content identity and availability are separate checks: a correct reference can identify an artifact the executor cannot fetch. Test the authorized path under actual execution conditions, including identity, network and destination. Do not use laptop credentials as proof of runner capability. Artifact Registry cleanup should respect the recovery options the team decided to retain. If a version matches deletion and retention policies, retention wins. Dry run allows selection to be observed before enabling deletion; use the required Data Access write logs and allow the periodic execution to occur. An empty query immediately after configuring the policy does not establish correct retention. In the exercise, the inventory says image A is available and configuration 7 is unknown. This is an evidence gap. Do not replace configuration with latest to obtain a green result. If a missing artifact must be rebuilt, treat the result as a candidate to validate: the same commit does not guarantee that every dependency and build tool is identical to those used previously.

3. Migrate data with explicit contracts

Consider A reading account_ref and B reading account_key. Renaming the column while A serves requests breaks its contract. One possible strategy adds the new representation, defines how writes remain coherent, backfills historical data, reconciles and changes readers under control. Removing the old representation is a later decision, conditional on ending dependencies and the approved return path. This design requires its own tests; adding a column does not automatically solve coherence or operational impact. In one rehearsal, a backfill finishes at 22:00 and old writes continue for seven minutes. Matching row counts do not establish matching values. Define how to identify concurrent changes, repeat batches without unwanted effects and verify the complete population. A useful sample must not be presented as complete reconciliation. Include infrequent consumers. A monthly batch still querying account_ref prevents removing the column merely because the frontend now uses B. In PostgreSQL 18, consult the specific ALTER TABLE subcommand to identify the required lock. Logical query compatibility does not eliminate lock waits. Plan duration, a waiting limit and contention observation before choosing the execution window.

4. Assess rollback against present state

Rolling back a Kubernetes Deployment restores an earlier revision's Pod template. An external database migration and completed effects in other services are not automatically reversed by that operation. Similarly, Cloud Deploy creates a new rollout based on an earlier release; a functional option must be selected for the present context. Yesterday's success label is not a compatibility test against today's data. Imagine B removed a field and committed new operations. Restoring A can cause further failures if A depends on that field. Recreating an empty column may make startup pass while leaving financial results incorrect. The decision should consider containing new writes, restoring a compatible contract, correcting the running version and reconciling effects. These are engineering options to evaluate using evidence, not automatic instructions for every incident. The test matrix should represent transition states: A with the new schema, B with historical data, coexisting writers and recovery after B produces new formats. Record results by combination. B being able to read v1 does not establish that A can read v2. This asymmetry explains why an apparently compatible upgrade can remove the possibility of a simple rollback.

5. Include configuration and secret validity

Reproducing code does not necessarily reproduce execution context. If configuration uses latest to fetch a secret, two executions of the same image can request different versions. Referencing a numeric version connects configuration, rehearsal and change. Store the reference in recovery records; avoid including the secret value in tickets, logs or teaching examples. Several conditions are distinct. The version must exist, be in a state permitting access and be accessible to the correct identity. A disabled version is not automatically replaced by the next enabled version when the request specifies its number. The stored credential must also remain valid in the accepting system. If a partner revoked it, successfully reading it from Secret Manager does not restore its validity. In an APS case, image A uses secret 7 and image B uses 8. Before proposing A for recovery, ask whether rotation retained an authorized coexistence period and how partner authentication was tested. If the old secret was compromised, the plan needs a secure combination with valid credentials. Pressure to return quickly to older code does not justify reintroducing the known cause of an exposure.

6. Handle messages that outlive deployment

A queue contains application history. Returning to an earlier producer changes future emissions but does not automatically change messages already published. If B introduced a state A cannot interpret, the backlog can remain incompatible after producer recovery. Inventory stored formats and formats that still-active writers can emit. Include retries and periodic flows that may reappear after the window. For Avro messages in Pub/Sub, the writer schema revision matters to reading. The googclient_schemarevisionid attribute helps identify that revision; parser strategy should account for schema resolution. An unknown revision requires diagnosis, not an assumption that all bytes follow the newest contract. The exercise's v1 and v2 tokens are application-contract labels, not an emulation of Pub/Sub compatibility mechanisms. Structural validity does not establish correct meaning either. An integer amount can remain valid when it changes from cents to euros. Use tests with expected outcomes for units, enumerations, missing values and business rules. Define handling for incompatible items while preserving traceability and avoiding duplicate effects. A consumer that stops failing because it acknowledges messages without processing them can reduce backlog while increasing functional impact.

7. Recover the complete inference flow

A recovery model includes more than stored weights. Identify the explicit version, preprocessing, feature order and units, serving code, available capacity and consumers of decisions. An alias can change versions; record the concrete identifiers used in rehearsal. In the consulted references, some Vertex AI learning-path pages redirect to Gemini Enterprise Agent Platform documentation; the consultation date does not represent a confirmed new exam edition. In the original example, M1 receives [age, balance] and M2 receives [balance, age]. Changing only the model can produce successful HTTP responses with incorrect meaning. Validate the combination using known inputs and use-case criteria. A model present in the registry does not establish ready serving with contingency capacity. The plan must include resources and time needed to make that option usable. If the model routed 180 critical incidents to low priority, restoration improves future decisions but does not automatically correct earlier ones. Define the affected population, reassessment, owner and operational acceptance. An authorized manual mode may be needed during containment. Distinguish decision quality, endpoint availability and reconciliation of effects; these are different outcomes that should appear in the recovery report.

8. Exercise: compare contracts with evidence

The local Python program compares three contract families: schema, message and feature. The candidate declares labels it accepts. Evidence declares stored labels and those that may be emitted during recovery. Names are teaching conventions: a schema label is not SQL analysis, and feature does not execute a model. Lists must come from actual inventory and tests outside this program. For each family, the algorithm finds observed labels the candidate does not accept. It retains the stored or emitters origin to guide action. None means unknown; an empty list means absence has been established. The program also receives explicit availability for required artifacts. False produces a blocker, None a gap, and True is only user-supplied evidence. A known conflict remains visible even when another field is unknown. Run the supplied examples. Then add v2 to backlog, remove v2 from candidate capabilities and observe the conflict. Repeat with unknown emitters and explain why the result can no longer be settled. Finally, declare every contract compatible: eligible-for-review still does not authorize production. The program does not test payloads, capacity, locks, credentials or cloud behavior. Deliver a decision record stating evidence limits and the operational checks still needed.

"""Original offline exercise: declared recovery contracts, not cloud validation.

Tokens are explicit application-contract labels, not provider schema IDs.
None means unknown; [] means a known empty population. No inferred compatibility.
"""
import copy
import hashlib
import itertools
import json
from pathlib import Path

KINDS = ('schema', 'message', 'feature')

def labels(value, nullable=False):
 if value is None and nullable:
 return None
 if not isinstance(value, list) or any(type(x) is not str or not x.strip for x in value):
 raise ValueError('expected a list of nonempty contract labels')
 if len(value)!= len(set(value)):
 raise ValueError('duplicate contract label')
 return set(value)

def evaluate(candidate, evidence):
 """Check reader compatibility with stored and possible future contracts.

 Artifact state is supplied, not discovered. False and unknown stay distinct.
 Every candidate required artifact needs an explicit state; extra states fail.
 Unknown reader capability preserves inventory gaps rather than inventing them.
 """
 if type(candidate) is not dict or set(candidate)!= {'accepts', 'artifacts'}:
 raise ValueError('candidate fields')
 if type(evidence) is not dict or set(evidence)!= {'stored', 'emitters', 'artifacts'}:
 raise ValueError('evidence fields')
 for mapping in (candidate['accepts'], evidence['stored'], evidence['emitters']):
 if type(mapping) is not dict or set(mapping)!= set(KINDS):
 raise ValueError('all three contract kinds are required')
 artifacts = labels(candidate['artifacts'])
 if not artifacts:
 raise ValueError('at least one required artifact')
 states = evidence['artifacts']
 if type(states) is not dict or set(states)!= artifacts:
 raise ValueError('artifact evidence must match requirements exactly')
 if any(v is not None and type(v) is not bool for v in states.values):
 raise ValueError('artifact state must be boolean or None')
 conflicts, unknown = [], []
 for kind in KINDS:
 accepted = labels(candidate['accepts'][kind], True)
 if accepted is None:
 unknown.append({'kind': kind, 'origin': 'candidate'})
 for origin in ('stored', 'emitters'):
 actual = labels(evidence[origin][kind], True)
 if actual is None:
 unknown.append({'kind': kind, 'origin': origin})
 elif accepted is not None:
 for token in sorted(actual - accepted):
 conflicts.append({'kind': kind, 'origin': origin, 'contract': token})
 unavailable = sorted(k for k, v in states.items if v is False)
 unknown_artifacts = sorted(k for k, v in states.items if v is None)
 blocked = bool(conflicts or unavailable)
 incomplete = bool(unknown or unknown_artifacts)
 return {
 'status': 'blocked' if blocked else 'incomplete' if incomplete else 'eligible-for-review',
 'conflicts': conflicts, 'unknownContracts': unknown,
 'unavailableArtifacts': unavailable, 'unknownArtifacts': unknown_artifacts,
 'evidenceComplete': not incomplete,
 'eligibleForReview': not blocked and not incomplete,
 'productionAuthorized': False,
 }

def baseline:
 return (
 {'accepts': {k: ['v1'] for k in KINDS}, 'artifacts': ['image:a', 'config:7']},
 {'stored': {k: ['v1'] for k in KINDS},
 'emitters': {k: ['v1'] for k in KINDS},
 'artifacts': {'image:a': True, 'config:7': True}},
 )

def exercise:
 fixtures = []
 def record(name, edit, expected):
 c, e = baseline; edit(c, e)
 before = copy.deepcopy((c, e)); out = evaluate(c, e)
 assert (c, e) == before
 assert out['status'] == expected, name
 fixtures.append({'id': name, **out})
 record('compatible', lambda c, e: None, 'eligible-for-review')
 record('stored-incompatible', lambda c, e: e['stored'].update(message=['v2']), 'blocked')
 record('future-incompatible', lambda c, e: e['emitters'].update(message=['v2']), 'blocked')
 record('unknown-emitter', lambda c, e: e['emitters'].update(message=None), 'incomplete')
 record('known-empty', lambda c, e: (e['stored'].update(message=[]), e['emitters'].update(message=[])), 'eligible-for-review')
 record('unknown-reader', lambda c, e: c['accepts'].update(feature=None), 'incomplete')
 record('missing-artifact', lambda c, e: e['artifacts'].update({'image:a': False}), 'blocked')
 record('unknown-artifact', lambda c, e: e['artifacts'].update({'image:a': None}), 'incomplete')
 record('conflict-and-gap', lambda c, e: (e['stored'].update(schema=['s3']), e['artifacts'].update({'config:7': None})), 'blocked')
 record('new-reader-accepts-old', lambda c, e: c['accepts'].update(message=['v1', 'v2']), 'eligible-for-review')
 record('both-origins-conflict', lambda c, e: (e['stored'].update(message=['v2']), e['emitters'].update(message=['v2'])), 'blocked')
 contract_combinations = 0
 # Three observations per kind: supported, unsupported, unknown.
 # Independent oracle computes known conflicts and evidence gaps over six cells.
 for values in itertools.product((['v1'], ['v2'], None), repeat=6):
 c, e = baseline
 for (origin, kind), value in zip(itertools.product(('stored', 'emitters'), KINDS), values):
 e[origin][kind] = value
 out = evaluate(c, e)
 assert len(out['conflicts']) == sum(v == ['v2'] for v in values)
 assert len(out['unknownContracts']) == sum(v is None for v in values)
 assert out['eligibleForReview'] == all(v == ['v1'] for v in values)
 contract_combinations += 1
 artifact_combinations = 0
 for values in itertools.product((True, False, None), repeat=2):
 c, e = baseline; e['artifacts'] = dict(zip(c['artifacts'], values))
 out = evaluate(c, e)
 assert len(out['unavailableArtifacts']) == sum(v is False for v in values)
 assert len(out['unknownArtifacts']) == sum(v is None for v in values)
 assert out['eligibleForReview'] == all(v is True for v in values)
 artifact_combinations += 1
 # Input label order and irrelevant dictionary order do not affect decisions.
 c, e = baseline; c['accepts']['message'] = ['v1', 'v2']
 e['stored']['message'] = ['v1', 'v3', 'v4']; expected = evaluate(c, e)
 order_permutations = 0
 for order in itertools.permutations(e['stored']['message']):
 e['stored']['message'] = list(order); assert evaluate(c, e) == expected
 order_permutations += 1
 invalid = [
 lambda c, e: c.update(extra=True),
 lambda c, e: e.update(extra=True),
 lambda c, e: c.pop('accepts'),
 lambda c, e: e['stored'].pop('schema'),
 lambda c, e: e['emitters'].update(extra=[]),
 lambda c, e: c['accepts'].update(message='v1'),
 lambda c, e: e['stored'].update(message=['v1', 'v1']),
 lambda c, e: e['emitters'].update(message=['']),
 lambda c, e: e['stored'].update(message=[1]),
 lambda c, e: e['artifacts'].update({'image:a': 1}),
 lambda c, e: e['artifacts'].update({'image:a': 'true'}),
 lambda c, e: e['artifacts'].pop('image:a'),
 lambda c, e: e['artifacts'].update(extra=True),
 lambda c, e: c.update(artifacts=[]),
 lambda c, e: c.update(artifacts=['image:a', 'image:a']),
 lambda c, e: c['accepts'].update(schema=[' ']),
 ]
 for mutate in invalid:
 c, e = baseline; mutate(c, e)
 try:
 evaluate(c, e)
 except ValueError:
 pass
 else:
 raise AssertionError('invalid input accepted')
 return {'scriptSha256': hashlib.sha256(Path(__file__).read_bytes).hexdigest,
 'fixtures': fixtures, 'contractCombinations': contract_combinations,
 'artifactCombinations': artifact_combinations, 'orderPermutations': order_permutations,
 'invalidInputs': len(invalid), 'inputPreserved': True,
 'cloudExecuted': False, 'databaseExecuted': False,
 'network': False, 'persistentWrites': False}

if __name__ == '__main__':
 print(json.dumps(exercise, ensure_ascii=False, sort_keys=True, indent=2))
IN PRACTICE

Image A is available but only reads v1 and v2 messages remain unprocessed. Recovery must address that contract even when rollout is technically executable.

Common pitfalls

Treating earlier success as current compatibility; forgetting batches and historical messages; using latest as identity; confusing a readable secret with a valid credential.

Related topics: Data migration and reconciliation · Secret and artifact management · Service and ML model recovery

Take this idea with you

A recovery option is a demonstrable combination of contracts and dependencies. Record conflicts and gaps without turning a local exercise into operational authorization.

Create account

Reference: Professional Cloud DevOps Engineer exam guide · Current linked guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.