← Professional Data Engineer: pipelines and data decisions
16 / 21 · 135 MIN

Operate retention and recovery sets

Connect retention, versions and dependencies to a recoverable set; distinguish metadata, estimates and evidence of service recovery.

1. Define the state that must return

In a fictional APS example, an error changes movements used for close and the team looks for the newest backup. However, calculation also depends on reference prices and classification rules. Recovering movements alone can produce an inconsistent result. Start with the assets the service needs, their relationships and the business point to recover. These examples do not represent BNP Paribas procedures. Distinguish availability, retention and recovery. A replica can keep the service available while reproducing an incorrect change. A copy can exist but be inaccessible because a key or permission is missing. A restore can finish before the consumer can use the data. Define RPO for tolerated data loss and RTO for recovery time, with observable acceptance criteria. Record incident time, the interval in which corruption could have started and the evidence supporting that boundary. A timestamp written in an inventory does not prove transactional consistency. Relate the chosen point to application execution and reconciliation. Where uncertainty exists, retain it in the decision and identify the missing check before using the copy.

2. Operate Bigtable versions and backups

A garbage-collection policy identifies cells eligible for removal, but cleanup runs in the background. If a consumer must not read expired versions, apply filters consistent with the policy. Do not conclude that configuration failed merely because an eligible cell remains visible. For combined criteria, read the rule as a deletion condition: intersection requires all criteria; union requires at least one. For example, to remove a version only when it is over 30 days old and a newer version exists, combine age and maximum-one-version rules with intersection. An age-only rule can remove the sole old version; a version-only rule can remove recent history. Predict both counterexamples before changing policy and consider existing data. When restoring a backup, prepare the new table and its policies: garbage collection is not inherited. Confirm the CMEK version pinned to the backup because key rotation does not change that association. The old version must be enabled for decryption. In replicated instances, copy time does not prove inclusion of every write from another cluster. Validate the captured point and consumer behavior after recovery.

3. Recover the warehouse with the right scope

In BigQuery, a deleted table uses the time-travel window in effect at deletion. Increasing the window from two to seven days afterwards does not retroactively extend that recovery. Treat current configuration and deletion history as separate evidence. Beyond time travel, fail-safe is not an extension directly queryable with SQL; emergency recovery requires Cloud Customer Care. Do not promise a self-service deadline without support for that commitment. A snapshot also has a concrete scope. Streaming-buffer data is not included. Confirm the captured set before deleting or changing the source. Time travel recovers data but does not restore table metadata; compare configuration and protection separately. A correct monetary total does not prove that a recovered table is ready for every consumer. When shortening partition expiration, consider existing partitions: those already beyond the new window expire immediately. Expiration and asynchronous physical deletion differ. A monthly consumer may depend on partitions ignored by the daily benchmark. Before the change, connect intended savings to historical use, owner acceptance and the recovery path actually rehearsed.

4. Restore objects without repeating effects

Soft delete retains deleted resources during the applicable window but does not keep them as live objects for ordinary reads. A change from seven to 30 days affects deletions after it takes effect; already deleted versions retain their original duration. Do not use current policy as proof that any old object remains recoverable. Identify the version and its specific condition. The restoration interface also matters. Restoring a bucket through CLI or JSON API restores the empty bucket; objects require separate recovery. Restoring an object creates a new live version and can trigger OBJECT_FINALIZE consumers. A pipeline that sends notifications or starts processing needs to control those effects. The same name does not prove the business operation has been recognized as a repeat. For concurrent operations, use appropriate preconditions. If you read generation G1 and another writer creates G2, deletion by name alone can affect the wrong version. A generation-match precondition binds the request to the observed version. Treat a failed precondition as new information requiring reassessment; blindly retrying without it removes the protection you intended.

5. Run local copy selection

The Python exercise below receives fictional metadata; it reads no backups and performs no restores. All times are integers on a common minute scale. The envelope defines asOf, incidentAt, targetRpo, targetRto, validationMinutes, requiredAssets and copies. There are one to 20 asset identities and one to 100 copies. Each copy has a unique id, asset, checkpoint, expiresAt, restoreMinutes, readable and keyAvailable. The model requires a checkpoint strictly before the incident and an incidentAt-minus-checkpoint age within inclusive RPO. The source must expire strictly after asOf plus restoreMinutes. readable and keyAvailable accept true, false or null; only true makes the dependency eligible. False and unknown appear as different reasons without being converted into success. The program groups by checkpoint and requires every asset at that same declared point. For each asset it chooses the fastest eligible copy, using id as a deterministic tie-breaker. It assumes parallel copies without contention followed by validation: total duration is elapsed time since the incident plus the maximum copy time plus validationMinutes. The RTO target is inclusive. Among sets meeting these rules, it selects the latest checkpoint. These teaching bounds do not represent cloud service contracts.

6. Predict results and challenge the model

Run python3 content/labs/pde-restore-set/run.py < content/labs/pde-restore-set/case.json from the project root. In the fixture, asOf=120 and incidentAt=100. Assets are ledger and prices. ledger95 exists but prices95 does not. Pair 90 has an unknown prices key. Pair 80 has true evidence, copy times of 6 and 7, and validation of 5. With RPO30 and RTO35, the model selects L80 and P80 and estimates 32 minutes from the incident, of which 12 remain. Set P90's key to true and confirm pair 90 is selected with a 34-minute total estimate. Then reduce RTO to 32: set 80 becomes the alternative fitting the model again. Test the expiration boundary: P80 with expiresAt=127 fails because 120+7 equals expiration; 128 passes. Compare this strict boundary with acceptance of RPO and RTO exactly at their limits. A common checkpoint does not prove byte consistency. Times are supplied assumptions without measurement of network, contention or startup. rtoProven, consistencyProven, cloudRestorePerformed and productionApproval remain false. Explain which rehearsal and reconciliation would replace each uncertainty with evidence. Do not change a flag to manufacture that proof.

7. Interpret quality and governance after recovery

A recovered platform needs more than readable tables. Confirm catalog references, owners, consumer access and policies supporting use. The exam guide uses Dataplex terminology; current inspected documentation presents auto data quality under Knowledge Catalog. Keep that mapping explicit without inventing a new exam version from a product-name change. A scan evaluates only configured rules and population. Incremental filters and sampling can exclude restored history. A green result over recent data does not prove preservation of six months of reporting. Record filters, execution date, rules, thresholds and the population actually considered. That description allows the data owner to judge whether evidence suits the intended use. A 95% threshold allows a rule to pass with 970 valid rows out of 1000; 30 exceptions still exist. Do not confuse aggregate passing with correctness of every row. Investigate exception relevance and confirm that no critical rule was omitted. For relationships across assets, add reconciliation of keys, grain and values rather than relying only on row counts.

8. Deliver a verifiable recovery procedure

Prepare a procedure another team can execute and evaluate. Identify assets, copy versions, recovery point, key and access dependencies, isolated destination, reconciliation criteria and return-to-service decision. Separate copy completion from functional acceptance. Measure each stage during rehearsal and identify dependencies on authorization or external support. Include stop conditions: an expired copy, unknown key, checkpoint missing a required asset, incompatible reference data or a monthly consumer outside test scope. An older copy can be usable within objectives, but the decision needs known loss and acceptance. A newer copy does not compensate for a missing essential asset. As a final exercise, submit a note for sets 80 and 90: explain initial selection, evidence that would permit changing the choice and checks after restoration. Add a retention change and assess its impact on consumers and recovery. The summary is to identify a complete set, confirm dependencies, recover within controlled scope, reconcile and obtain service acceptance. Connect these stages to ingestion, contracts and capacity in preceding lessons.

"""Original parallel-copy planning model; metadata is supplied, never measured."""
import json
import re
import sys


def identifier(value):
 return isinstance(value, str) and re.fullmatch(r'[A-Za-z0-9_-]{1,64}', value, re.ASCII) is not None


def integer(value):
 return type(value) is int and 0 <= value <= 10**9


def evaluate(data):
 keys = {'asOf', 'incidentAt', 'targetRpo', 'targetRto', 'validationMinutes', 'requiredAssets', 'copies'}
 if not isinstance(data, dict) or set(data)!= keys:
 raise ValueError('invalid planning envelope')
 for key in keys - {'requiredAssets', 'copies'}:
 if not integer(data[key]):
 raise ValueError('invalid time or duration')
 if data['incidentAt'] > data['asOf']:
 raise ValueError('incident cannot be after assessment')
 assets = data['requiredAssets']
 if not isinstance(assets, list) or not 1 <= len(assets) <= 20 or not all(identifier(a) for a in assets) or len(set(assets))!= len(assets):
 raise ValueError('expected 1..20 unique assets')
 copies = data['copies']
 if not isinstance(copies, list) or not 1 <= len(copies) <= 100:
 raise ValueError('expected 1..100 copies')
 ids, assessed = set, []
 copy_keys = {'id', 'asset', 'checkpoint', 'expiresAt', 'restoreMinutes', 'readable', 'keyAvailable'}
 for copy in copies:
 if not isinstance(copy, dict) or set(copy)!= copy_keys:
 raise ValueError('invalid copy fields')
 if not identifier(copy['id']) or copy['id'] in ids or not identifier(copy['asset']) or copy['asset'] not in assets:
 raise ValueError('invalid copy identity or asset')
 ids.add(copy['id'])
 for key in ['checkpoint', 'expiresAt', 'restoreMinutes']:
 if not integer(copy[key]):
 raise ValueError('invalid copy time')
 if copy['checkpoint'] > data['asOf'] or copy['expiresAt'] < copy['checkpoint']:
 raise ValueError('invalid copy chronology')
 for key in ['readable', 'keyAvailable']:
 if copy[key] is not None and type(copy[key]) is not bool:
 raise ValueError('evidence must be true,false or null')
 reasons = []
 if copy['checkpoint'] >= data['incidentAt']:
 reasons.append('not-before-incident')
 if data['incidentAt'] - copy['checkpoint'] > data['targetRpo']:
 reasons.append('outside-rpo')
 if copy['expiresAt'] <= data['asOf'] + copy['restoreMinutes']:
 reasons.append('expires-before-copy-completes')
 for key in ['readable', 'keyAvailable']:
 if copy[key] is not True:
 reasons.append(key + ('-unknown' if copy[key] is None else '-false'))
 assessed.append({'id': copy['id'], 'asset': copy['asset'], 'checkpoint': copy['checkpoint'], 'eligible': not reasons, 'reasons': sorted(reasons)})
 assessment = {r['id']: r for r in assessed}
 plans = []
 for checkpoint in sorted({c['checkpoint'] for c in copies}, reverse=True):
 selected, missing = [], []
 for asset in sorted(assets):
 options = [c for c in copies if c['asset'] == asset and c['checkpoint'] == checkpoint and assessment[c['id']]['eligible']]
 options.sort(key=lambda c: (c['restoreMinutes'], c['id']))
 if options:
 selected.append(options[0])
 else:
 missing.append(asset)
 remaining = None if missing else max(c['restoreMinutes'] for c in selected) + data['validationMinutes']
 minutes = None if missing else data['asOf'] - data['incidentAt'] + remaining
 plans.append({'checkpoint': checkpoint, 'copyIds': [c['id'] for c in selected], 'missingEligibleAssets': missing, 'estimatedRemainingMinutes': remaining, 'estimatedMinutes': minutes, 'withinModel': not missing and minutes <= data['targetRto']})
 feasible = [p for p in plans if p['withinModel']]
 return {'selected': feasible[0] if feasible else None, 'plans': plans,
 'copies': sorted(assessed, key=lambda r: r['id']),
 'parallelCopyAssumption': True, 'rtoProven': False,
 'consistencyProven': False, 'cloudRestorePerformed': False,
 'productionApproval': False}


def unique_object(pairs):
 result = {}
 for key, value in pairs:
 if key in result:
 raise ValueError('duplicate JSON key')
 result[key] = value
 return result


if __name__ == '__main__':
 try:
 raw = sys.stdin.read(1000001)
 if len(raw) > 1000000:
 raise ValueError('input too large')
 print(json.dumps(evaluate(json.loads(raw, object_pairs_hook=unique_object)), sort_keys=True))
 except (ValueError, TypeError, RecursionError) as error:
 print('invalid input: ' + str(error), file=sys.stderr)
 sys.exit(2)
IN PRACTICE

Pair80 is selected when prices90 has an unknown key; the 32-minute estimate does not prove real RTO.

Common pitfalls

Latest copy per asset as a coherent set; TTL as immediate exclusion; current policy as retroactive retention; partial scans as full validation.

Related topics: Fidelity and partitions · Rejections and replay · Capacity and continuity

Take this idea with you

A copy existing does not prove recovery: confirm the set, dependencies, data and measured time.

Create account

Reference: Professional Data Engineer standard exam guide · Current linked standard guide (document title v4.2); edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.