← Professional Cloud DevOps Engineer: delivery and reliability
23 / 25 · 135 MIN

Reliability: dependencies, restore and recovery evidence

Prepare workload and data recovery, calculate dependencies and distinguish estimates from evidence for service reopening.

1. Define what recovery means

A completed restore is a technical event. Deciding whether a service can be used again requires a recovery contract: covered users, required operations, acceptable data state and required evidence. In a fictional funds exercise, the landing page might work while instruction processing remains unavailable. If the approved criterion includes looking up an instruction and checking its document, opening only the landing page does not satisfy that criterion. Write the acceptance operation before the drill, with its expected result and a person responsible for evaluation. RTO and RPO measure different dimensions. The former bounds the period of unavailability defined in the agreement; the latter bounds the tolerated time window of data loss. Identify exactly where measurement starts and stops. In our exercise, interruption starts at 14:00, the restore command starts at 14:08, finishes at 14:24 and business validation ends at 14:33. If the agreement ends at that validation, the observation is 33 minutes. Reporting 16 minutes measures only the command. Before promising a time, distinguish objective, estimate and observation. The objective guides design; the estimate depends on assumptions; the observation documents an executed drill. Keep these three values separately. A 30-minute forecast does not yet demonstrate that the team recovered the service in that time.

2. Inventory the recoverable set

Start with one user operation and identify what it needs to finish. In the APS example, a Pod receives an instruction, queries Cloud SQL, reads a document and records the outcome. GKE workload backup can protect manifests and volumes covered by the service, but does not automatically include external Cloud SQL state or image layers. The target cluster must also exist with the Backup for GKE agent enabled. These boundaries require planning for components that do not appear as tasks within the same backup. Create a working table with component, recovery mechanism, required identity, dependency and restore evidence. For the image, record its digest and where the content will remain available. An intact manifest reference is unhelpful if the artifact was removed and no target node has a cache. For the external document, identify who demonstrates recovery and how it relates to the instruction. For the cluster, check networking configuration and capacity through the process that actually creates it. The table’s purpose is to find omissions before the intervention window. Do not fill an unknown cell with “included in backup” for convenience. Assign an owner to the gap and rehearse the required path. Project work includes these dependencies even when they belong to different teams.

3. Check keys as recovery dependencies

An accessible copy and a usable key are distinct pieces of evidence. In Backup for GKE, a volume already protected with CMEK can retain a dependency on the original disk key even when the plan uses another key. Storing the backup in another region therefore does not alone demonstrate independence from the original region. In the fictional drill, the team can list the copy but volume restoration fails when using the key. The ticket should identify that specific dependency, the identity attempting access and the observed condition, without exposing secret material. In Cloud KMS, a DISABLED version still has key material and can be re-enabled with authorization. The version must be usable for the required cryptographic operation. Creating a different primary version does not replace the version protecting existing content. Avoid turning time pressure into a destruction operation: removing a key does not retroactively re-encrypt data. In the exercise plan, “keys ready” is a task with its own evidence. Check the authorized path with representative identities and retain the drill result. If only the usual administrator has confirmed access, access by the actual recovery identity remains unproven. The timing estimate should declare this condition. A key blocker with no known resolution time cannot remain hidden inside a task marked as three minutes.

4. Select and validate coherent state

Cloud SQL PostgreSQL PITR creates a new instance. Plan validation of that destination and connection transition instead of assuming the source changed state. Also confirm the instance’s recoverable window. If the proven window ends at 10:42, there is no basis for promising 10:47 through that path. This limits demonstrated recovery; it does not prove that every later record was destroyed. Further investigation may be possible, but communication should separate possibility from evidence. Service recovery requires consistency across components. In our case, the restored table contains a reference to a document created at 09:58, while documents were recovered to 09:55. APIs can respond normally while the reference remains invalid. Define a check crossing that relationship: select instructions from the affected interval, resolve associated documents and record exceptions. Equal counts do not establish matching identities. Treat the former destination as part of the plan too. In self-managed PostgreSQL after promotion, the former primary must be prevented from returning to accept writes as though it remained authoritative. The exercise does not prescribe commands for a managed service; it teaches investigation of the risk of two histories. A forgotten monthly job can keep writing to the former destination after the interactive application switches. Include it in reconciliation and confirmation of the transition.

5. Calculate dependencies and timing margin

The local model represents tasks with an identifier, duration in minutes and a list of predecessors. It assumes enough execution resources, failure-free tasks and each task starting as soon as all predecessors finish. These assumptions permit a planning calculation; they do not simulate operator queues, approval requests or cloud service availability. If two tasks need the same person, add the constraint to the plan or use another model. Do not declare parallelism merely because two lines appear alongside each other. In the example, the cluster takes 12 minutes, the database 18, control of the former writer five and keys three; they can start at time zero. Workload needs the cluster and keys, then takes eight minutes. It finishes at minute 20. Validation needs workload, database and writer control; it takes six minutes and finishes at 26. Reopening takes another four, finishing at 30. Adding every duration incorrectly counts parallel work. Using only the longest individual duration ignores successor tasks. Now increase cluster preparation to 17 minutes. Workload finishes at 25 and reopening at 35. The 30-minute target no longer fits the model. Conversely, shortening the key task from three minutes to one does not change the initial deadline because the cluster still dominates that dependency. Choose improvements based on the chain that actually limits the outcome.

6. Separate false conditions from unknown ones

The exercise requests five evidence fields: coherent data, former writer controlled, keys accessible, functional check and business acceptance. Each field accepts true, false or null. These names and this list are a teaching convention for the exercise, not a mandatory Google Cloud standard or an internal BNP Paribas procedure. In a real implementation, criteria would need to match the service, risk and approved responsibilities. Supplying true is a user-provided assertion; the program does not inspect the evidence supporting it. A false condition represents a failed check. Null indicates missing confirmation. Both prevent the worksheet from being ready for a decision, but they require different actions: remedy a known failure or collect missing evidence. The age of the recovered point can also be unknown. Do not replace it with the age of the newest backup while data consistency remains under investigation. The readyForDecision result appears only when the estimate fits RTO, reported age fits RPO and all five fields are true. Even then, productionAuthorized stays false. The model has no operational authority, verifies no signatures and does not establish current conditions. It organizes a discussion that can be checked. The responsible person must examine evidence, understand exceptions and decide under the process applicable to the service.

7. Communicate partial recovery and learn from the drill

An incident update should make remaining impact and the next decision understandable. In the scenario, the internal probe passes but partners still fail through an older path. Communicate partial recovery, affected users, ongoing actions and the next update time. Avoid both declaring full success and discarding useful probe evidence. Its result is valid within the path it actually exercised. Record observation limits alongside the conclusion so the next team does not interpret it as universal coverage. During coordination handover, confirm that someone has accepted responsibility, knows which actions are running and understands pending reopening conditions. A list of links without context does not show which decisions have already been made. In our exercise, restore is running and the key is a known blocker; the successor should know who is investigating the key, which operations cannot proceed and when the forecast will be reconsidered. Business communication may be in English, with explicit times and scope. After the drill, connect failures to trackable actions. For the key dependency, propose an owner, deadline and another drill where that dependency is unavailable. The expected outcome should be demonstrable. A completed presentation or a generic request for greater care does not establish that the recovery path has been repaired.

8. Run and interpret the local exercise

Run python3 run.py in a local directory using Python 3.13 or compatible. The program uses only the standard library and calls no APIs, restores no databases and changes no files. Its JSON output is a model report using fictional data. Start with parallel-plan: the estimated result is 30 minutes and the worksheet is ready for a decision with the supplied values. Then compare rto-missed, rpo-missed and unknown-recovery-point. A time that fits the plan does not compensate for data loss outside the limit or a coherent point that remains unknown. In cleanup-outside-service, a cleanup task finishes after reopening. The program lists it under outsideServicePath, retaining 30 minutes for the explicitly selected service endpoint task. This does not decide whether excluding cleanup from the real contract is correct. If cleanup is an acceptance condition, change the dependencies and calculate again. The user is responsible for the graph’s meaning. The code checks eight named cases, 512 duration combinations and 243 combinations of evidence fields. It rejects 14 invalid inputs, including cycles, repeated tasks and undefined references. Validation also confirms that input data was not modified. These checks demonstrate properties of this small model. They do not establish cloud timing, backup consistency, effective permissions or actual failover capability. Use the following questions to explain decisions and limitations as well as reproduce numbers.

"""Original recovery planning worksheet. No scheduler, database or cloud execution."""
from copy import deepcopy
from graphlib import TopologicalSorter, CycleError
import hashlib
import itertools
import json
from pathlib import Path

GATES = ('data_consistent', 'old_writer_fenced', 'keys_accessible',
 'functional_check', 'business_acceptance')


def minutes(value, name):
 if type(value) is not int or value < 0:
 raise ValueError(name + ' must be a nonnegative integer')
 return value


def assess(tasks, service_task, rto_minutes, recovery_age_minutes, rpo_minutes, checks):
 """All tasks start as soon as dependencies finish, with unlimited workers.

 Durations are fictional estimates from interruption time zero. Checks are
 supplied assertions, not verified evidence. None means unknown. Tasks outside
 the service-task ancestor graph are explicitly listed. RPO uses an externally
 established coherent recovery point; a timestamp alone does not establish one.
 """
 for name, value in [('rto', rto_minutes), ('rpo', rpo_minutes)]:
 minutes(value, name)
 if recovery_age_minutes is not None:
 minutes(recovery_age_minutes, 'recovery age')
 if not isinstance(tasks, list) or not tasks:
 raise ValueError('tasks must be a nonempty list')
 if not isinstance(checks, dict) or set(checks)!= set(GATES):
 raise ValueError('provide exactly the five evidence fields')
 if any(value is not None and type(value) is not bool for value in checks.values):
 raise ValueError('checks must be true, false or null')
 by_id = {}
 for task in tasks:
 if not isinstance(task, dict) or set(task)!= {'id', 'minutes', 'after'}:
 raise ValueError('invalid task fields')
 name = task['id']
 if not isinstance(name, str) or not name.strip or name in by_id:
 raise ValueError('task IDs must be nonempty and unique')
 minutes(task['minutes'], name)
 deps = task['after']
 if not isinstance(deps, list) or any(not isinstance(x, str) for x in deps):
 raise ValueError('dependencies must be strings in a list')
 if len(set(deps))!= len(deps):
 raise ValueError('repeated dependency')
 by_id[name] = task
 if not isinstance(service_task, str) or service_task not in by_id:
 raise ValueError('service task not defined')
 graph = {name: task['after'][:] for name, task in by_id.items}
 if any(dep not in by_id for deps in graph.values for dep in deps):
 raise ValueError('dependency not defined')
 try:
 order = list(TopologicalSorter(graph).static_order)
 except CycleError as error:
 raise ValueError('cyclic recovery plan') from error
 finish, ancestors = {}, {}
 for name in order:
 finish[name] = max((finish[d] for d in graph[name]), default=0) + by_id[name]['minutes']
 ancestors[name] = {name}.union(*(ancestors[d] for d in graph[name]))
 total = finish[service_task]
 rpo_met = None if recovery_age_minutes is None else recovery_age_minutes <= rpo_minutes
 return {
 'estimatedServiceMinutes': total,
 'finishMinutes': dict(sorted(finish.items)),
 'outsideServicePath': sorted(set(graph) - ancestors[service_task]),
 'estimatedRtoMet': total <= rto_minutes,
 'reportedRpoMet': rpo_met,
 'failedChecks': sorted(k for k, v in checks.items if v is False),
 'unknownChecks': sorted(k for k, v in checks.items if v is None),
 'readyForDecision': total <= rto_minutes and rpo_met is True and all(v is True for v in checks.values),
 'productionAuthorized': False,
 }


def task(name, duration, *after):
 return {'id': name, 'minutes': duration, 'after': list(after)}


def evidence:
 tasks = [task('cluster', 12), task('keys', 3), task('database', 18),
 task('fence', 5), task('workload', 8, 'cluster', 'keys'),
 task('validate', 6, 'database', 'workload', 'fence'),
 task('service', 4, 'validate')]
 good = dict.fromkeys(GATES, True)
 base = [tasks, 'service', 30, 4, 5, good]
 original = deepcopy(base)
 fixtures = []
 for name, change in [
 ('parallel-plan', {}), ('rto-missed', {2: 29}),
 ('rpo-missed', {3: 6}), ('unknown-recovery-point', {3: None}),
 ('unknown-acceptance', {5: {**good, 'business_acceptance': None}}),
 ('old-writer-active', {5: {**good, 'old_writer_fenced': False}}),
 ('cleanup-outside-service', {0: tasks + [task('cleanup', 60, 'service')]}),
 ('slow-cluster', {0: [{**t, 'minutes': 17} if t['id'] == 'cluster' else t for t in tasks]}),
 ]:
 args = deepcopy(base)
 for index, value in change.items:
 args[index] = value
 fixtures.append({'id': name, **assess(*args)})
 assert fixtures[0]['estimatedServiceMinutes'] == 30
 assert fixtures[0]['readyForDecision'] is True
 assert all(not f['productionAuthorized'] for f in fixtures)
 assert all(not f['readyForDecision'] for f in fixtures[1:6])
 assert fixtures[6]['estimatedServiceMinutes'] == 30
 assert fixtures[6]['outsideServicePath'] == ['cleanup']
 assert fixtures[7]['estimatedServiceMinutes'] == 35
 duration_checks = 0
 for a, b, c in itertools.product(range(8), repeat=3):
 model = [task('a', a), task('b', b), task('c', c, 'a', 'b')]
 result = assess(model, 'c', 100, 0, 0, good)
 assert result['estimatedServiceMinutes'] == max(a, b) + c
 assert assess(list(reversed(model)), 'c', 100, 0, 0, good) == result
 duration_checks += 1
 gate_checks = 0
 for values in itertools.product((True, False, None), repeat=5):
 checks = dict(zip(GATES, values))
 result = assess(tasks, 'service', 30, 4, 5, checks)
 assert result['readyForDecision'] == all(v is True for v in values)
 assert result['unknownChecks'] == sorted(k for k,v in checks.items if v is None)
 assert result['failedChecks'] == sorted(k for k,v in checks.items if v is False)
 gate_checks += 1
 invalid = [
 {0: []}, {0: tasks + [tasks[0]]}, {0: [task('a', -1)]},
 {0: [task('a', True)]}, {0: [task('service', 1, 'missing')]},
 {0: [task('service', 1, 'x'), task('x', 1, 'service')]},
 {0: [task('service', 1, 'service')]}, {1: 'absent'},
 {2: True}, {3: -1}, {4: 1.2}, {5: {}},
 {5: {**good, 'functional_check': 1}},
 {0: [task('a', 1), task('service', 2, 'a', 'a')]},
 ]
 for change in invalid:
 args = deepcopy(base)
 for index, value in change.items:
 args[index] = value
 try:
 assess(*args)
 except ValueError:
 pass
 else:
 raise AssertionError('invalid input accepted')
 assert base == original
 return {'fixtures': fixtures, 'durationCombinations': duration_checks,
 'gateCombinations': gate_checks, 'invalidInputs': len(invalid),
 'inputPreserved': True, 'cloudExecuted': False, 'databaseRestored': False,
 'network': False, 'persistentWrites': False, 'independentVerification': False,
 'scriptSha256': hashlib.sha256(Path(__file__).read_bytes).hexdigest}


if __name__ == '__main__':
 print(json.dumps(evidence, sort_keys=True, indent=2))
IN PRACTICE

Cluster 12 minutes, workload 8, validation 6 and reopening 4 form a 30-minute path, with database recovery in parallel.

Common pitfalls

Accessible backup as restore proof; timestamps as consistency; estimates as measurements; null as approval; an internal probe as coverage of every client.

Related topics: Processing continuity and reconciliation · Identity and key management · Incident communication and shift handover

Take this idea with you

A defensible recovery connects executable dependencies, coherent data, write control and demonstrated acceptance, with explicit evidence limits.

Create account

Reference: Disaster recovery planning guide · Current linked guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.