← Professional Cloud DevOps Engineer: delivery and reliability
13 / 15 · 135 MIN

SRE: measurement, capacity and recovery

Define SLIs, interpret budgets and coordinate capacity, recovery and service-retirement decisions.

1. Write the measurement contract

A fictional team supports an API for checking processing-instruction status. Before choosing a percentage, write down who uses the API, which operations count, when waiting starts and what constitutes a useful outcome. This exercise counts valid real requests, including requests that fail inside the service. Synthetic tests and invalid requests are excluded under a rule approved before measurement. A processing failure does not turn a valid request into an excluded request. Store a versioned definition with an owner, window and data origin so APS and development can reproduce the report. In a batch of 1,200 records, excluding 200 tests and 100 distinct invalid requests leaves 900 eligible requests. If 18 failed, the success indicator is 882/900, or 98%. For overnight processing, choose a unit representing delivery: unique valid files available by the agreed deadline. If 95 of 100 arrived by 06:00 and five only at 06:12, the on-time result is 95%, even if every process exits with 0. These numbers and rules are original teaching examples unrelated to any bank’s policies. At handover, ask the functional owner to confirm the unit and deadline before discussing alerts.

2. Calculate without losing the population

Two intervals contain 90 successful requests out of 100 and 891 out of 900. Adding counts gives 981/1,000, or 98.1%. The simple mean of 90% and 99% gives 94.5% because it weights intervals with different volumes equally. A windows-based indicator answers another question: how many windows met their defined condition? Do not switch between those units within one report. In Cloud Monitoring, inspect the effective SLO definition and distinguish requestBased from windowsBased before interpreting the dashboard. Now define an internal indicator where each good request must succeed and finish within 250 ms. Among 10,000 requests, 80 failed, 150 were slow and 30 belong to both groups. There are 80 + 150 − 30 = 200 bad requests. Subtract the overlap once because each request occupies only one position in the denominator. This composite calculation is a teaching model, not a ready-to-apply BasicSli configuration. Before implementation, identify metrics or events that can represent the intersection. Two independent totals without overlap information cannot reconstruct this result exactly. Request an event sample and check cases at the 250 ms boundary.

3. Handle gaps and govern the budget

The ingress register identifies 400 eligible requests, but the collector retains only 360 outcomes, all successful. The other 40 have unknown outcomes. Reporting 100% for the entire population hides the collection failure; reporting 90% as proven failure also exceeds the evidence. Record incomplete coverage, investigate missing identifiers and apply only an agreed unknown-data rule. A conservative policy may exist, but it must be explicit. In meetings, communicate what happened to the service separately from what can be measured. The budget also has units. With 800,000 requests and a 99.95% target, 400 bad requests are allowed. There is no unique conversion to minutes without knowing affected traffic in each interval. Under a rolling 28-day period, repairing today does not erase yesterday’s impact. If policy includes supplier failures, excluding them retroactively to unblock a release changes the agreement. Keep the calculation and refer the decision to whoever can authorize exceptions. Policies should define owners, actions and exceptions; Google’s published example is neither a universal obligation nor your organization’s internal policy.

4. Exercise: reconcile before concluding

Save this lesson’s code as run.py and execute python3 run.py with Python 3.13. The exercise uses only the standard library and prints JSON. Its fictional population contains 900 eligible requests, 200 synthetic records and 100 invalid records. Among eligible requests, 18 fail, 36 exceed 250 ms and nine belong to both groups. Predict the results first: 45 bad, 855 good, an SLI of 19/20 and a budget of nine under the teaching target of 99%. Remaining budget is −36 requests; do not round it up to zero, which would hide the size of the excess. Then read missing and unexpected. Removing outcome r899 or introducing an unknown identifier leaves coverage unreconciled; sli and slo_met become null. An empty population also proves no success. Compare latency_boundary, budget_equality and fractional_budget to understand inclusion of 250 ms, exact equality with the target and a budget smaller than one request. The program also explores 24 classification combinations and rejects seven invalid inputs. These trials check the local model. The expected inventory is supplied as ground truth for the exercise: there is no real ingress query, log-integrity verification or cloud call. Production use would also require establishing how that inventory is obtained and protected.

5. Validate capacity before failover

The team plans to move 18,000 requests/s into a region tested to sustain 12,000/s within the agreed latency. Without immediate expansion, at least 6,000/s of capacity is missing. Decide in advance which traffic may wait and who authorizes priority. Limit admission to sustainable capacity and monitor queue age and completed work. An unbounded queue merely transfers the problem to memory, deadlines and later recovery. Include the time needed to clear accumulated work after restoration. Available quota permits allocation within limits but does not guarantee physical capacity. A VM reservation in another zone does not automatically cover the chosen zone; check matching zone and properties. In Cloud Run, a single-threaded application with four vCPUs can saturate one thread while average CPU remains near 25%. Review concurrency, latency and CPU together instead of assuming useful spare capacity. For a Kubernetes HPA using only CPU utilization percentage, missing effective requests prevent the relevant utilization from being defined. Check requests and HPA conditions before changing replica limits. Each mechanism requires its own evidence; the load trial still needs representative traffic and dependencies.

6. Contain amplification and finish work

Retrying an operation may help with a transient failure, but three layers with three total attempts each can generate 27 backend calls for one initial request if all fail and exhaust their attempts. Draw where each retry occurs and define a global budget alongside waiting and spreading attempts. Also establish whether the operation can be repeated without duplicating effects. In the lesson scenario, load already exceeds capacity: immediately retrying every rejected request increases pressure on the same dependency. Propagate the remaining deadline. If a read has an 800 ms total budget and has already spent 620 ms, at most 180 ms remain, ignoring transport margins here. Restarting 800 ms at every service allows work to continue after the client gives up. When removing a VM from an unmanaged instance group used by an Application Load Balancer, coordinate draining with the process lifecycle. Configuring 90 seconds at the load balancer cannot keep a process alive if your automation terminates it after five. Observe in-flight requests and termination criteria; reused connections and other load-balancer types have details requiring the actual configuration to be checked. This exercise neither executes draining nor establishes zero request loss.

7. Demonstrate recovery and prepare retirement

In a fictional recovery trial, the incident starts at 11:00, recovered data reaches 10:42 and functional service returns at 11:30. With a 15-minute RPO and a 45-minute RTO, the 30-minute restoration meets its objective, but the recovered point loses 18 minutes and misses RPO. Record both results. A reachable endpoint alone does not establish usable data: add reconciliation, business operations and dependencies to acceptance. Retain timestamps and criteria so another team can reproduce the reasoning. After migration, seven days without traffic do not resolve uncertainty about a monthly consumer. Confirm that consumer’s schedule, owner and destination before retiring the source. Review backups, recovery paths and dependencies outside migration scope as well. Prepare a decision with evidence, owners and an observation period appropriate to the service. Update runbooks, diagrams, monitoring and maintenance procedures following the decision. Cost reduction is sustainable only if retirement preserves still-needed capabilities; include decommissioning and validation work in the plan instead of leaving an ownerless task after delivery.

8. Coordinate the decision and hand over to RUN

In the final case, one region fails, remaining capacity is limited and part of the telemetry disappears. One team wants to divert traffic while another prepares a return to the previous route. Before incompatible changes execute, establish single coordination, an operations owner and someone responsible for communication. Record the hypothesis, authorized action, expected outcome and reassessment time. Limiting traffic under approved priorities can provide useful mitigation while capacity recovers. The absence of errors in the few observed outcomes does not justify announcing full recovery. Prepare a shift handover that lets a colleague continue without reconstructing the incident. Include measured capacity, admitted and deferred requests, backlog age, telemetry coverage, the latest change and the next decision criterion. Practise a concrete business update: “Service is partially restored; lower-priority work remains deferred and telemetry reconciliation is still in progress.” Give the next update time without promising a recovery time unsupported by evidence. The lesson summary is a working sequence: define the population, reconcile outcomes, calculate in the correct unit, limit load, demonstrate recovery and confirm dependencies before retirement. Reuse these criteria in alerting, FinOps and production-acceptance topics.

"""Original SLI teaching model: fixed window, fictional ingress inventory."""
from fractions import Fraction
from hashlib import sha256
from itertools import product
from pathlib import Path
import json
import platform


def assess(records, expected_eligible_ids, target=Fraction(99, 100)):
 if not isinstance(target, Fraction) or not 0 < target < 1:
 raise ValueError('target must be an exact fraction strictly between zero and one')
 seen, eligible = set, []
 for row in records:
 if not {'id', 'synthetic', 'valid', 'success', 'duration_ms'} <= row.keys:
 raise ValueError('missing field')
 if not isinstance(row['id'], str) or not row['id'] or row['id'] in seen:
 raise ValueError('empty or duplicate event identifier')
 if any(type(row[k]) is not bool for k in ['synthetic', 'valid', 'success']):
 raise ValueError('classification flags must be booleans')
 if type(row['duration_ms']) is not int or row['duration_ms'] < 0:
 raise ValueError('duration must be a nonnegative integer')
 seen.add(row['id'])
 if not row['synthetic'] and row['valid']:
 eligible.append(row)
 ids = {r['id'] for r in eligible}
 expected = set(expected_eligible_ids)
 complete = ids == expected
 good = sum(r['success'] and r['duration_ms'] <= 250 for r in eligible)
 failed = sum(not r['success'] for r in eligible)
 slow = sum(r['duration_ms'] > 250 for r in eligible)
 both = sum(not r['success'] and r['duration_ms'] > 250 for r in eligible)
 bad = len(eligible)-good
 assert bad == failed+slow-both
 evaluable = complete and bool(expected)
 budget = len(eligible)*(1-target) if evaluable else None
 return {'records': len(records), 'eligible': len(eligible), 'good': good,
 'bad': bad, 'failed': failed, 'slow': slow, 'both': both,
 'missing': sorted(expected-ids), 'unexpected': sorted(ids-expected),
 'coverage_complete': complete,
 'sli': str(Fraction(good, len(eligible))) if evaluable else None,
 'allowed_bad': str(budget) if evaluable else None,
 'remaining_bad': str(budget-bad) if evaluable else None,
 'slo_met': bad <= budget if evaluable else None}


def event(name, synthetic=False, valid=True, success=True, duration=250):
 return dict(id=name, synthetic=synthetic, valid=valid, success=success, duration_ms=duration)


def run:
 expected = {f'r{i}' for i in range(900)}
 records = [event(f'r{i}', success=i >= 18, duration=400 if i < 9 or 18 <= i < 45 else 250) for i in range(900)]
 records += [event(f's{i}', synthetic=True) for i in range(200)]
 records += [event(f'i{i}', valid=False) for i in range(100)]
 complete = assess(records, expected)
 assert (complete['eligible'], complete['good'], complete['bad']) == (900, 855, 45)
 assert (complete['failed'], complete['slow'], complete['both']) == (18, 36, 9)
 assert complete['sli'] == '19/20' and complete['allowed_bad'] == '9'
 assert complete['remaining_bad'] == '-36' and complete['slo_met'] is False
 missing = assess([r for r in records if r['id']!= 'r899'], expected)
 assert missing['sli'] is None and missing['missing'] == ['r899'] and missing['slo_met'] is None
 extra = assess(records+[event('unknown')], expected)
 assert extra['sli'] is None and extra['unexpected'] == ['unknown']
 empty = assess([], set)
 assert empty['sli'] is None and empty['slo_met'] is None
 boundary = assess([event('at-250'), event('at-251', duration=251)], {'at-250','at-251'})
 assert boundary['good'] == 1 and boundary['bad'] == 1
 equality = assess([event(f'e{i}', success=i > 0) for i in range(100)], {f'e{i}' for i in range(100)})
 assert equality['remaining_bad'] == '0' and equality['slo_met'] is True
 fractional = assess([event('tiny')], {'tiny'})
 assert fractional['allowed_bad'] == '1/100'
 assert Fraction(90+891, 100+900) == Fraction(981,1000)
 assert 80+150-30 == 200
 assert Fraction(800000)*(1-Fraction(9995,10000)) == 400
 assert 3**3 == 27 and 800-620 == 180
 assert 18000-12000 == 6000
 truth_cases = 0
 for synthetic, valid, success, duration in product([False,True], [False,True], [False,True], [249,250,251]):
 row = event('t', synthetic, valid, success, duration)
 chosen = not synthetic and valid
 result = assess([row], {'t'} if chosen else set)
 assert result['eligible'] == int(chosen)
 assert result['good'] == int(chosen and success and duration <= 250)
 truth_cases += 1
 malformed = [event('bad', duration=-1),event('bad', duration=True),event('bad', duration=1.5),event('bad', success='true'),event(''),{'id':'missing'}]
 invalid = 0
 for row in malformed:
 try:
 assess([row], {'bad'})
 except ValueError:
 invalid += 1
 try:
 assess([event('duplicate'),event('duplicate')], {'duplicate'})
 except ValueError:
 invalid += 1
 assert invalid == 7
 return {'scriptSha256':sha256(Path(__file__).read_bytes).hexdigest,
 'pythonVersion':platform.python_version,
 'fixtures':{'complete':complete,'missing':missing,'unexpected':extra,'empty':empty,'latency_boundary':boundary,'budget_equality':equality,'fractional_budget':fractional},
 'truthCases':truth_cases,'invalidInputCases':invalid,
 'calculationChecks':6,'network':False,'persistentWrites':False,
 'vendorExecution':False,'independentVerification':False,
 'limitations':'Fixed-window trusted synthetic inventory; no actual ingress reconciliation, histogram query, collector, cloud API, or production acceptance decision.'}


if __name__ == '__main__':
 print(json.dumps(run, ensure_ascii=False, indent=2))
IN PRACTICE

A failover leaves 6,000 requests/s above capacity and incomplete collection prevents declaring full recovery.

Common pitfalls

Excluding failures from the denominator, presenting unknown outcomes as success, confusing quota with capacity and retiring monthly dependencies.

Related topics: Alerting and telemetry coverage · FinOps and resilient-capacity cost · Production acceptance and application retirement

Take this idea with you

A reliability decision needs a known population, demonstrated capacity and checkable recovery outcomes.

Create account

Reference: Implementing SLOs · Current linked guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.