← Professional Cloud DevOps Engineer: delivery and reliability
25 / 25 · 135 MIN

Reliability economics and recovery capacity

Compare contingency options against recovery objectives, capacity and equivalent costs, distinguishing quota, reservations, commitments and missing evidence.

1. Define the service the budget must recover

A contingency option should be compared against the business need it intends to satisfy. In a fictional funds-processing service, starting a VM is only one step. The service may need coherent data, capacity for accumulated demand, partner authentication and functional validation before resuming. This example does not represent internal BNP Paribas standards. Monetary amounts in this lesson are fictional and are not Google Cloud prices. Start by recording objectives: maximum recovery time, acceptable data point, minimum capacity and required functions. If a degraded mode exists, identify who accepts it, which operations it permits and for how long. An option serving queries but not processing reconciliation must not be compared as if it recovered the whole service. Then build an evidence record for each alternative. Include rehearsal conditions, observed duration, recovered-point age, supported load, dependencies and costs over the same horizon. An unknown value must remain unknown. The purpose is an explicit decision about cost and delivered service. A saving that removes a required capability changes risk; present that effect alongside the reduction in the monthly bill.

2. Distinguish quota,reservation and actual consumption

Quota is a permitted usage limit. It does not guarantee the requested resource exists in a zone when it needs to be created. A recovery plan presenting only free quota leaves availability unproven. Identify the location, machine type and resources required, then check the capacity mechanism planned for the scenario. A Compute Engine reservation has matching rules. The zone and relevant VM properties must match; a reservation in zone A is not automatically consumable in zone B. For a specifically targeted reservation, affinity must also select the correct reservation. Do not treat vCPUs from different configurations as freely interchangeable slots. A rehearsal using the actual template helps uncover incompatibilities before the window. Confirm capacity still available. If a reservation has 20 slots and 18 remain occupied by services that cannot stop, two slots are free for new VMs. Counting the 18 already used again overstates the contingency option. If the plan depends on releasing capacity, include the sequence, impact and authorization for that release. Record observed state and date; a reservation existing today can have different consumption when an incident occurs.

3. Separate financial discount from recovery capacity

A usage commitment and a reservation should be analyzed by their respective functions. The former is a financial decision subject to scope and conditions; the latter addresses capacity with consumption requirements. Relationships can exist between them for particular resources, but they are not equivalent concepts. Consult product terms rather than automatically attributing a technical guarantee to a discount. In an original budget, the team reduces estimated cost because it anticipates a commitment. It still needs to demonstrate capacity at the recovery location, compatible resources and a usable activation path. Likewise, a reservation no longer being consumed should not be treated as cost-free capacity. Comparison must identify obligations remaining when workload moves regions. To discuss the proposal with FinOps, present expected utilization, variability, horizon and alternative scenarios. Distinguish cost avoidable now from already committed cost and incremental contingency activation cost. Avoid counting the same benefit in two categories. If the architecture decision changes, review financial applicability and technical capacity separately. A clear estimate can reveal that the cheapest option during normal operation is not cheapest across the planned drills and activations.

4. Size demand after failure

Normal and contingency capacity can differ substantially. In this lesson's model, two regions each receive 60 operations per second. Each region was tested up to 100. If one fails and all its demand moves to the other, the survivor receives 120 against capacity of 100. Capacity in the unavailable region cannot remain included in the plan. Evaluate alternatives against the required service: demonstrated additional capacity, priority for critical functions, limiting nonessential load or an approved degraded mode. Adding retries indiscriminately can increase demand without producing useful outcomes. If backlog exists, supporting the normal arrival rate alone is insufficient; a plan must also recover accumulated work within the relevant deadline. Use a representative rehearsal to measure the limiting resource. More workers may not increase throughput if the constraint is database connections, storage, an external API or lock contention. Record operation mix, data size, concurrency and quality criteria. Capacity of 120 used in the exercise is supplied input, not a measurement performed by the program. Revalidate it when demand characteristics or a relevant dependency change.

5. Evaluate Spot against deadlines and tolerance

Spot VMs use excess capacity whose availability varies and can be interrupted. The discount can help workloads that tolerate interruption and resume under control. It does not turn Spot into a capacity guarantee for a service with a strict deadline. The design should explain what happens when capacity is unavailable, a VM ends during processing or remaining time no longer permits another attempt. In a fictional exercise, preparation tasks can use checkpoints and idempotent effects while the critical recovery path has a demonstrated capacity alternative. That separation must be tested. Storing checkpoints only on a disk that disappears with the worker can make retry policy ineffective. Record where restart state lives and how much work might be repeated. Compare costs by scenario, including repeated work, activation time and operational maintenance. A low unit price can coexist with higher total cost when interruptions are frequent. Do not assign the platform a universal preemption frequency: use declared assumptions and sensitivity analysis. The decision should connect workload tolerance with the business deadline and evidence available for the alternative when Spot is unusable.

6. Compare costs over a common horizon

First define the accounting scope: standby, replication, retention, transfer, applicable licenses, observability, preparation and drill execution. Some amounts are continuous; others arise per activation. Also identify costs common to alternatives and excluded costs. A table containing only VMs can be useful as one component but must not be presented as total service cost. We use a deliberately simple annual model: cost=12×monthly cost+activation cost×activation count. Option A costs 100 units per month and 300 per activation; B costs 150 and 100. With no activations, A totals 1200 and B 1800. With four, A totals 2400 and B 2200. They tie with three activations. These invented values explore a decision; they do not describe an invoice or actual prices. Activation count can represent planned drills and incidents included in a scenario, but is not automatically a forecast of failure frequency. Present multiple scenarios when that frequency is uncertain. Do not add incremental costs already included in the monthly amount. When material components are missing, preserve the gap and avoid artificial precision. Use comparison to guide evidence collection and discussion with service and budget owners.

7. Connect savings to capabilities being removed

A lightly used environment may exist for recovery. Before decommissioning it, check rollback-plan references, data dependencies and replacement conditions. If the return option still depends on it, deletion is more than resource cleanup. It may be a valid change, but needs a demonstrated alternative or an explicit decision about risk and service objectives. Also examine the benefit attributed to redundancy. Two regions depending on the same key unavailable in the selected scenario can fail together. Duplicating compute does not resolve that dependency. A more expensive design should be connected to the failure modes it can actually cover. Do not attribute independence to components merely because they have different names or locations. In the committee presentation, show cost, delivered capacity, rehearsal results and gaps. If A costs less but supports 90 operations per second when 120 are needed, place that deviation alongside the financial difference. Distinguish an option that failed a requirement from another lacking evidence. The owner should understand what they are approving, the conditions under which the comparison holds and when analysis needs to be repeated.

8. Exercise: filter constraints before ranking prices

The local program receives options containing monthly and per-activation costs in fictional cents, observed recovery duration, recovered-point age, capacity and declared resource availability. It compares values against supplied objectives. None means unknown. A known failed requirement remains visible even when another field is missing. The program neither confirms rehearsals nor queries quotas, reservations or prices. Ranking includes only options satisfying declared constraints with calculable cost in the scenario. If two options tie, both appear in the lowest-cost set. Input order does not decide a winner. The result concerns only options demonstrated by supplied data; it does not prove an unknown option could never become better. Even with all inputs complete, production still receives no automatic authorization. Run the four-activation scenario first, then zero and three. Reduce the cheaper option's capacity to 90 and observe its exclusion from ranking. Replace capacity with None and compare known failure with incomplete evidence. Finally, leave activation cost unknown: with zero activations that field does not affect total cost; with one it prevents calculation. Explain each result and identify operational evidence still needed to turn the comparison worksheet into an executable proposal.

"""Original offline recovery comparison. All prices and measurements are fictional.

This is a supplied-evidence model, not a cloud quote or production authorization.
Annual cost = 12 * monthly fixed cents + activation count * incremental cents.
"""
import copy
import hashlib
import itertools
import json
from pathlib import Path

FIELDS = {'id','monthlyCents','activationCents','recoveryMinutes','recoveryPointAgeMinutes','capacity','resourcesAvailable'}

def integer(value, nullable=False):
 if value is None and nullable: return
 if type(value) is not int or value < 0: raise ValueError('nonnegative integer required')

def evaluate(options, activations, max_rto, max_rpo, min_capacity):
 for value in (activations,max_rto,max_rpo,min_capacity): integer(value)
 if type(options) is not list or not options: raise ValueError('nonempty option list required')
 ids=set;result=[]
 for option in options:
 if type(option) is not dict or set(option)!=FIELDS: raise ValueError('option fields')
 ident=option['id']
 if type(ident) is not str or not ident.strip or ident in ids: raise ValueError('unique nonempty ID required')
 ids.add(ident)
 for key in FIELDS-{'id','resourcesAvailable'}: integer(option[key],True)
 if option['resourcesAvailable'] is not None and type(option['resourcesAvailable']) is not bool:
 raise ValueError('resourcesAvailable must be boolean or None')
 failed=[];unknown=[]
 for key,limit,direction in [('recoveryMinutes',max_rto,'max'),('recoveryPointAgeMinutes',max_rpo,'max'),('capacity',min_capacity,'min')]:
 value=option[key]
 if value is None: unknown.append(key)
 elif (value>limit if direction=='max' else value<limit): failed.append(key)
 if option['resourcesAvailable'] is None: unknown.append('resourcesAvailable')
 elif option['resourcesAvailable'] is False: failed.append('resourcesAvailable')
 monthly,activation=option['monthlyCents'],option['activationCents']
 missing_cost=[]
 if monthly is None: missing_cost.append('monthlyCents')
 if activations and activation is None: missing_cost.append('activationCents')
 annual=None if missing_cost else 12*monthly+(activations*activation if activations else 0)
 status='fails-declared-constraints' if failed else 'evidence-incomplete' if unknown else 'meets-declared-constraints'
 result.append({'id':ident,'status':status,'failedConstraints':failed,'unknownConstraints':unknown,
 'missingCostFields':missing_cost,'annualCents':annual,
 'rankable':not failed and not unknown and annual is not None})
 result.sort(key=lambda r:r['id'])
 ranking=sorted(({'id':r['id'],'annualCents':r['annualCents']} for r in result if r['rankable']),
 key=lambda r:(r['annualCents'],r['id']))
 best=[r['id'] for r in ranking if r['annualCents']==ranking[0]['annualCents']] if ranking else []
 return {'options':result,'ranking':ranking,'lowestCostAmongDemonstratedOptions':best,
 'allInputsComplete':not any(r['unknownConstraints'] or r['missingCostFields'] for r in result),
 'realPricesVerified':False,'actualRecoveryProven':False,'productionAuthorized':False}

def candidate(ident='warm'):
 return {'id':ident,'monthlyCents':10000,'activationCents':30000,'recoveryMinutes':20,
 'recoveryPointAgeMinutes':5,'capacity':120,'resourcesAvailable':True}

def exercise:
 fixtures=[]
 def record(label,opts,count=4):
 before=copy.deepcopy(opts);r=evaluate(opts,count,30,10,120);assert opts==before
 fixtures.append({'id':label,**r});return r
 a=candidate('a');b=candidate('b');b.update(monthlyCents=15000,activationCents=10000)
 assert record('four-activations',[a,b])['lowestCostAmongDemonstratedOptions']==['b']
 assert record('no-activations',[a,b],0)['lowestCostAmongDemonstratedOptions']==['a']
 r=record('three-activations-break-even',[a,b],3)
 assert r['lowestCostAmongDemonstratedOptions']==['a','b']
 assert all(x['annualCents']==210000 for x in r['ranking'])
 a=candidate('cheap');a.update(monthlyCents=100,capacity=90)
 assert record('cheap-capacity-fails',[a,candidate])['lowestCostAmongDemonstratedOptions']==['warm']
 a=candidate('unknown');a.update(capacity=None)
 assert record('unknown-capacity',[a])['ranking']==[]
 a=candidate;a.update(recoveryMinutes=40,resourcesAvailable=None)
 r=record('known-failure-and-gap',[a]);assert r['options'][0]['failedConstraints']==['recoveryMinutes'];assert not r['allInputsComplete']
 a=candidate;a.update(monthlyCents=None)
 assert record('unknown-cost',[a])['ranking']==[]
 a=candidate;a.update(activationCents=None)
 assert record('unused-unknown-activation-cost',[a],0)['ranking'][0]['annualCents']==120000
 assert record('required-unknown-activation-cost',[a],1)['ranking']==[]
 a=candidate('b');b=candidate('a')
 assert record('equal-cost-tie',[a,b])['lowestCostAmongDemonstratedOptions']==['a','b']
 a=candidate;a.update(resourcesAvailable=False)
 assert record('resources-unavailable',[a])['ranking']==[]
 a=candidate;a.update(recoveryMinutes=30,recoveryPointAgeMinutes=10,capacity=120)
 assert record('inclusive-boundaries',[a])['ranking']
 a=candidate;a.update(monthlyCents=0,activationCents=0)
 assert record('explicit-zero-cost',[a])['ranking'][0]['annualCents']==0
 combinations=0
 for rto,rpo,capacity,available in itertools.product((20,40,None),(5,15,None),(120,90,None),(True,False,None)):
 a=candidate;a.update(recoveryMinutes=rto,recoveryPointAgeMinutes=rpo,capacity=capacity,resourcesAvailable=available)
 r=evaluate([a],4,30,10,120)['options'][0]
 failures=sum([rto==40,rpo==15,capacity==90,available is False])
 unknown=sum(v is None for v in (rto,rpo,capacity,available))
 assert len(r['failedConstraints'])==failures and len(r['unknownConstraints'])==unknown
 assert r['rankable']==(not failures and not unknown)
 combinations+=1
 cost_combinations=0
 for monthly,activation,count in itertools.product((0,100,None),(0,200,None),(0,1,4)):
 a=candidate;a.update(monthlyCents=monthly,activationCents=activation)
 r=evaluate([a],count,30,10,120)['options'][0]
 known=monthly is not None and (count==0 or activation is not None)
 assert (r['annualCents'] is not None)==known
 if known:assert r['annualCents']==monthly*12+(count*activation if count else 0)
 cost_combinations+=1
 opts=[candidate('c'),candidate('a'),candidate('b')];expected=evaluate(opts,4,30,10,120);permutations=0
 for order in itertools.permutations(opts):
 assert evaluate(list(order),4,30,10,120)==expected;permutations+=1
 invalid=[lambda o:o.clear,lambda o:o.append(copy.deepcopy(o[0])),lambda o:o[0].update(id=''),
 lambda o:o[0].update(extra=1),lambda o:o[0].pop('capacity'),lambda o:o[0].update(capacity=-1),
 lambda o:o[0].update(capacity=True),lambda o:o[0].update(monthlyCents=1.5),
 lambda o:o[0].update(activationCents='100'),lambda o:o[0].update(resourcesAvailable=1)]
 for mutate in invalid:
 opts=[candidate];mutate(opts)
 try:evaluate(opts,4,30,10,120)
 except ValueError:pass
 else:raise AssertionError('invalid option accepted')
 settings=[(-1,30,10,120),(True,30,10,120),(4,-1,10,120),(4,30,None,120),(4,30,10,1.5)]
 for args in settings:
 try:evaluate([candidate],*args)
 except ValueError:pass
 else:raise AssertionError('invalid setting accepted')
 return {'scriptSha256':hashlib.sha256(Path(__file__).read_bytes).hexdigest,'fixtures':fixtures,
 'constraintCombinations':combinations,'costCombinations':cost_combinations,'orderPermutations':permutations,
 'invalidInputs':len(invalid)+len(settings),'inputPreserved':True,
 'cloudExecuted':False,'network':False,'persistentWrites':False}

if __name__=='__main__':
 print(json.dumps(exercise,ensure_ascii=False,sort_keys=True,indent=2))
IN PRACTICE

A costs less and recovers in 20 minutes but supports only 90 operations/s against a need for 120. B costs more and satisfies rehearsed constraints. Comparison must present that deviation.

Common pitfalls

Treating quota as reservation; double-counting consumed capacity; treating Spot as a deadline guarantee; omitting drills and activation costs; deleting the only rollback to lower the bill.

Related topics: Recovery planning and RTO/RPO · Capacity management and overload · FinOps and decommissioning

Take this idea with you

Compare cost among options demonstrating the required service and keep gaps explicit. The lowest estimate does not fix unmet technical requirements.

Create account

Reference: Professional Cloud DevOps Engineer exam guide · Current linked guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.