← Professional Data Engineer: pipelines and data decisions
19 / 21 · 135 MIN

Design recovery with verifiable dependencies

Assess failover mechanism, protection, regional dependencies and critical path before promising service recovery.

Define recovery the business accepts

A fictional funds technology team learns of a regional outage. Another region holds a data copy, but the positions report also depends on identities, keys, execution capacity, routines and scheduled queries. The sponsor asks when the service can be used again. The answer should identify the recovered outcome: query positions from a known date, reconcile movements and deliver the report to its authorized consumer. Saying only that the dataset exists does not answer the question. Draw that outcome’s dependencies with their technical owners. For each, record estimated duration, prerequisites, readiness evidence and the owner of any gaps. Include alternative access paths and the location of incident evidence because the usual console may be unavailable. If the agreed RTO is thirty minutes from the incident and ten have elapsed, twenty remain for recovery and validation. The clock does not restart when the team opens the plan. This example teaches planning; durations are illustrative and represent neither vendor guarantees nor BNP Paribas procedures. Agree which consumer action demonstrates recovery before measuring whether the target was met.

Choose the mechanism for the failure scenario

With BigQuery cross-region replication, the secondary replica is read-only. Documentation distinguishes that replication from managed disaster recovery during a total primary-region outage. In the former case, secondary promotion depends on the primary being available. An architecture needing to resume writes during regional loss must explicitly assess the applicable recovery mechanism, edition and reservation instead of inferring that capability from the existence of a replica. Under managed disaster recovery, hard failover can proceed while the primary is unavailable but does not wait for unreplicated data. Soft failover waits for synchronization and requires both regions to be available. The choice should connect continuity, potential loss and actual incident conditions. Prepare a table of the known replication point, later writes and affected consumers. Do not turn a replication point into proof that every delivery to external systems also completed. A report may have been sent before the outage and need separate reconciliation. The rehearsal should test the complete path, including recovery of the evidence needed for the decision. Record what remains unknown rather than filling gaps with an assumed zero-loss outcome.

Recover protection and access with data

A disabled Cloud KMS version retains key material but cannot be used while in that state. Restoring a version scheduled for destruction places it in disabled state; it does not immediately make it usable. The plan should identify the required version, the control authorizing its use and evidence of successful destination reads. Administrative console access does not prove that the service identity can decrypt data. For a BigQuery replica using CMEK, verify the replica’s regional key configuration. If the source dataset has default_kms_key, creation requires replica_kms_key suitable for the destination region. Column protection also needs mechanism-specific treatment: taxonomy-based policy-tag bindings are not automatically repaired by promotion. Current documentation separately addresses policies assigned directly to columns or data governance tags; do not generalize one rule to every type. If custom masking is used, include its UDF and location. During an incident, avoid resolving a specific failure by opening general access. Record the correction, test the consumer identity and retain audit evidence. A restored data copy must remain protected while it serves as the operational source.

Inventory resources a copy does not make usable

A routine can appear in the secondary region while still referencing a regional connection in the source. An external table can have replicated metadata yet depend on bucket objects that do not satisfy required location rules. Inventory these resources as executable dependencies, with read and transformation trials. The name appearing in a catalogue is only a metadata observation. Include performance behavior in the plan. For search indexes, documentation describes metadata replication and rebuilding index data in the promoted region. Do not promise identical latency merely because the definition was copied. For materialized views, verify where referenced tables reside; replication does not remove location requirements. When several failover reservations participate in one query, confirm the product’s common-destination conditions. In the funds example, a trial must execute the consumer query under its identity, call required routines and check delivery time. Retain a negative case too, such as an unavailable external dependency, to demonstrate how the service fails and who must act. Such evidence supports a usable recovery procedure instead of a resource inventory alone.

Run the local recovery graph

Run python3 run.py < case.json in the pde-recovery-path lab. Times are minutes on a teaching clock. Input contains incidentAt, asOf, targetRto, maxEvidenceAge, targets and nodes. Each node has an identifier, dependencies, duration and evidence marked confirmed, missing or failed. Missing evidence requires a null timestamp; in other states the timestamp cannot be in the future. Duration represents work still to be performed from assessment time, not the total duration of a partly completed task. The fixture has two branches. access takes five and precedes engine, which takes six. key takes two and precedes storage, which takes eight. reconcile takes four and waits for both branches; consumer takes three. The longest path is access,engine,reconcile,consumer, taking eighteen minutes. Ten minutes have already elapsed since the incident, making planned total time 28. Evidence for key is missing, so the result is evidence-blocked even though the timing fits the thirty-minute RTO. The program calculates planned times even when evidence is missing, but excludes the unknown effort needed to resolve that gap. Read the blocker list alongside the timing result.

Distinguish critical path from blocker set

The critical path identifies the branch determining duration in a model with sufficient parallelism. It does not necessarily include every dependency that can prevent recovery. key is on the shorter branch and still blocks consumer. The program traverses all target ancestors to gather missing, failed or stale evidence. A node unrelated to that target should not block its decision; if it is another required target, it receives its own assessment. Set key to confirmed with evidence at time 110: the forecast becomes plan-feasible. Increase key’s duration to four: the critical path changes to key,storage,reconcile,consumer and total time becomes 29. At three, the branches tie; the program chooses the lexicographically smaller path solely to make output reproducible. This does not imply greater business priority. Tests permute node lists and dependency order. Cycles, unknown dependencies and duplicate references are rejected. Do not remove a problematic node just to obtain a green forecast: that changes recovery scope and needs its own justification. Check whether evidence remains valid for the current configuration before changing its status in the fixture.

Rehearse timing, evidence and capacity

With all evidence confirmed, targetRto=28 accepts the 28-minute forecast; targetRto=27 does not. If assessment only occurs at time 113, the total rises to 31 without changing any durations. This exercise exposes time already consumed by detection, decisions and preparation. Evidence age also has a boundary: with asOf=110 and a maximum of twenty, an observation from 90 is accepted; one from 89 is stale. These are teaching rules to adapt to the real service contract. The calculation assumes parallel execution without contention. If access and key depend on the same unavailable person, their durations cannot be treated as simultaneous without further analysis. The model also measures neither RPO, cross-dataset consistency nor ability to decrypt data. rtoProven and dataConsistencyProven remain false. Use the exercise to challenge assumptions and prepare real trials with observed start and finish, correct identities and reconciliation. Record forecast, measured execution and consumer acceptance separately. A favorable earlier rehearsal does not replace evidence after a relevant architecture or permissions change. Keep the actual conditions of each rehearsal with its results.

Resume deliveries and prepare return

Promotion does not automatically redirect scheduled BigQuery queries to the new region. The procedure must recreate them in the applicable destination and verify identity, parameters and write target. Original-region job history also does not automatically appear in the secondary region. Prepare evidence collection before an incident and avoid concluding a job never existed merely because a regional query returns no rows. Before removing the source, confirm promotion completed, consumers execute at the destination and deliveries were reconciled. Plan return with data produced during recovery and the access controls valid in that phase. At the decision meeting, present confirmed capabilities, blockers, timing estimate and potential loss separately. If evidence for a key dependency is missing, explain why a favorable timing estimate is insufficient. The RUN owner should be able to reproduce diagnosis without depending on the project author. Finish the exercise with a short English update naming the next owner and the evidence needed to resume service. Preserve unresolved conditions in the handover instead of treating a successful control-plane operation as complete business recovery.

"""Original offline recovery dependency model. No infrastructure operations."""
import json
import re
import sys


def require(ok, message):
 if not ok:
 raise ValueError(message)


def integer(value, low, high):
 return type(value) is int and low <= value <= high


def identifier(value):
 return type(value) is str and re.fullmatch(r'[A-Za-z0-9_-]{1,48}', value) is not None


def keys(value, fields):
 require(type(value) is dict and set(value) == set(fields.split), 'Unexpected fields')


def evaluate(payload):
 keys(payload, 'incidentAt asOf targetRto maxEvidenceAge targets nodes')
 for name in ['incidentAt', 'asOf', 'targetRto', 'maxEvidenceAge']:
 require(integer(payload[name], 0, 1000000000), 'Invalid time')
 require(payload['asOf'] >= payload['incidentAt'], 'Assessment precedes incident')
 nodes = payload['nodes']
 require(type(nodes) is list and 1 <= len(nodes) <= 30, 'Need 1..30 nodes')
 registry = {}
 for node in nodes:
 keys(node, 'id dependencies minutes evidence evidenceAt')
 require(identifier(node['id']) and node['id'] not in registry, 'Invalid or duplicate node id')
 require(integer(node['minutes'], 1, 10000), 'Invalid duration')
 require(node['evidence'] in ['confirmed', 'missing', 'failed'], 'Invalid evidence status')
 if node['evidence'] == 'missing':
 require(node['evidenceAt'] is None, 'Missing evidence requires null time')
 else:
 require(integer(node['evidenceAt'], 0, payload['asOf']), 'Invalid evidence time')
 deps = node['dependencies']
 require(type(deps) is list and len(deps) <= 29 and all(identifier(v) for v in deps), 'Invalid dependencies')
 require(len(set(deps)) == len(deps) and node['id'] not in deps, 'Duplicate or self dependency')
 registry[node['id']] = node
 for node in nodes:
 require(all(v in registry for v in node['dependencies']), 'Unknown dependency')
 targets = payload['targets']
 require(type(targets) is list and 1 <= len(targets) <= 10, 'Need 1..10 targets')
 require(all(identifier(v) and v in registry for v in targets) and len(set(targets)) == len(targets), 'Invalid targets')
 visiting, calculated = set, {}

 def visit(name):
 if name in calculated:
 return calculated[name]
 require(name not in visiting, 'Dependency cycle')
 visiting.add(name)
 node = registry[name]
 parents = [visit(v) for v in sorted(node['dependencies'])]
 critical = min(parents, key=lambda p: (-p['remaining'], p['path'])) if parents else None
 ancestors = {name}
 for parent in parents:
 ancestors.update(parent['ancestors'])
 result = {'remaining': node['minutes'] + (critical['remaining'] if critical else 0),
 'path': (critical['path'] if critical else []) + [name], 'ancestors': ancestors}
 calculated[name] = result
 visiting.remove(name)
 return result

 # Validate every component, including nodes outside requested target closures.
 for name in sorted(registry):
 visit(name)
 results = []
 for target in sorted(targets):
 plan = calculated[target]
 blockers = []
 for name in sorted(plan['ancestors']):
 node = registry[name]
 reasons = []
 if node['evidence']!= 'confirmed':
 reasons.append(node['evidence'])
 if node['evidenceAt'] is not None and payload['asOf'] - node['evidenceAt'] > payload['maxEvidenceAge']:
 reasons.append('stale')
 if reasons:
 blockers.append({'id': name, 'reasons': reasons})
 total = payload['asOf'] - payload['incidentAt'] + plan['remaining']
 fits = total <= payload['targetRto']
 results.append({'id': target, 'plannedRemainingMinutes': plan['remaining'],
 'plannedMinutesFromIncident': total,
 'plannedFinishAt': payload['asOf'] + plan['remaining'],
 'criticalPath': plan['path'], 'evidenceBlockers': blockers,
 'timeFitsRto': fits,
 'decision': 'evidence-blocked' if blockers else 'plan-feasible' if fits else 'deadline-infeasible'})
 return {'targets': results, 'timeUnit': 'minutes', 'allTargetsPlanFeasible': all(t['decision'] == 'plan-feasible' for t in results),
 'recoveryPerformed': False, 'rtoProven': False,
 'dataConsistencyProven': False, 'productionApproval': False}


def main:
 raw = sys.stdin.read(1000001)
 require(len(raw) <= 1000000, 'Input too large')
 print(json.dumps(evaluate(json.loads(raw)), sort_keys=True))


if __name__ == '__main__':
 try:
 main
 except (ValueError, TypeError, RecursionError):
 print('Invalid recovery-path fixture', file=sys.stderr)
 sys.exit(2)
IN PRACTICE

The fixture predicts eighteen remaining minutes and 28 since the incident, but missing key evidence blocks consumer even though key is outside the critical path.

Common pitfalls

Confusing a replica with write recovery, counting only critical-path dependencies, restarting the RTO clock, assuming promotion redirects schedules, or treating metadata as proven execution.

Related topics: Retention and recovery · Attempt control · Migration contracts

Take this idea with you

A forecast is useful when it states dependencies, evidence, elapsed time and consumer acceptance conditions.

Create account

Reference: Professional Data Engineer standard exam guide · Current linked standard guide (document title v4.2); edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.