← Professional Cloud Architect: architecture and operations
16 / 25 · 120 MIN

Infrastructure during maintenance and recovery

Prepare capacity, volumes, network policies and managed dependencies for a change that RUN can recover and maintain.

Separate capacity, eligibility and health

Before a maintenance window, map an instance’s replacement path: where it can run, which resources it needs, what storage it mounts and how it starts receiving traffic. An existing node is not automatically a usable destination. It may lack residual memory, sit outside the disk’s zone or fail workload selection constraints. A toleration permits a matching taint but does not exclusively select that pool; combine it with a selection constraint when that is the requirement. In this lesson’s fictional reporting case, the requirement is capacity for six Pods after one zone fails. That statement needs a different calculation and rehearsal from demonstrating six Ready Pods before maintenance. Record residual capacity, per-Pod requirements, failure domains and eligibility conditions. Use measured values and identify who updates them, because another workload can consume the margin between project design and change execution.

Give each control its proper scope

A PDB cannot prevent hardware failures and does not control a Deployment’s RollingUpdate. Workload strategy, disruption budget and node upgrades are different mechanisms. In a GKE surge upgrade, the additional node needs resources before replacing the old one. If quota permits only the four existing nodes, maxSurge=1 does not create a fifth slot through intent alone. Prepare that margin or rehearse and approve an alternative compatible with required availability. Do not assume a PDB blocks the operation indefinitely: the documented surge strategy honors the PDB and termination grace period for up to one hour, after which eviction can be forced. Confirm the specific strategy and applicable documentation before the change. Maintenance exclusions also have exceptions, including critical-vulnerability interventions. Communicate planned maintenance and recovery capability to the business without turning operational settings into an absolute guarantee of no disruption.

Preserve storage constraints

A volume can outlive a Pod yet remain unusable by the chosen replacement. Early zonal provisioning may place a disk somewhere incompatible with the consumer’s mandatory location. For new PVCs and compatible drivers, WaitForFirstConsumer accounts for Pod topology during binding and provisioning. It does not move an existing disk and should not be combined with nodeName as a shortcut because that field bypasses the scheduler. Another boundary concerns exclusivity: ReadWriteOnce is node-scoped and can allow multiple Pods on that node. Evaluate ReadWriteOncePod when single-Pod scope is required, confirming CSI support and application requirements. In decommissioning plans, separate technical removal from actually releasing costs. An already locked Bucket Lock policy cannot be shortened; creating a bucket for future data does not release old objects. Include these constraints in budget, change sequence and handover before assuming that removing an application removes all associated resources.

Validate the complete connectivity path

Investigate each layer from evidence. An accepted PSC endpoint establishes connection state but does not establish a custom hostname in the application’s resolver. If a test with the correct IP and name works while the application receives NXDOMAIN, check the private record and zone visibility. With enforced native NetworkPolicy, an egress-isolated source and ingress-isolated destination must both permit the flow. Permissions from multiple policies combine; adding a restrictive policy does not remove an existing broad allow. Confirm the effective set rather than only the file just changed. Public NAT provides outbound connectivity and associated reply traffic but does not publish a service for new incoming connections. In rehearsal, record source, destination, name, port, workload identity and observed result. A test from an administrator’s laptop does not replace the actual service path.

Exercise: calculate slots without pooling incompatible resources

The local model takes each node’s residual budget in millicores and MiB. For identical Pods, it computes how many copies fit in CPU and memory and uses the smaller integer count. On a node with 6000 millicores and 6144 MiB, a Pod needing 2000 millicores and 4096 MiB fits only once: memory limits the result. Two nodes with 1500 free millicores each also cannot host a Pod requesting 2000; CPU does not pool across nodes to execute one Pod. Predict those two outcomes before running the code. Run python3 content/labs/pca-infrastructure-capacity/run.py and also observe the effect of marking a node ineligible. The program validates integer units, unique IDs and known excluded nodes. It neither contacts Kubernetes nor predicts the real scheduler: it omits volumes, affinity, topology spread, quotas and concurrency. A fitsModel=true result is only a capacity condition in this simplified model, never production approval.

Compare one-node maintenance with zone loss

The main case has four nodes, two per zone, each with 4000 residual millicores and 8192 MiB. Every Pod requests 2000 millicores and 3072 MiB. Each node provides two slots, eight in total. Removing one node leaves six slots; removing both nodes in one zone leaves four. The requirement for six replicas after zone failure is unmet even when one-node maintenance passes the calculation. voluntary_allowance shows only simplified healthy-minus-minimum arithmetic; it does not reproduce the controller’s disruptionsAllowed or prevent involuntary failure. To correct the design, consider additional eligible capacity or a formally accepted degraded mode with lower load. Then rehearse latency and actual behavior. Also test database failover: Cloud SQL HA can recover on the same address while the client retains closed connections. The application must restore access and handle interrupted operations according to safe retry rules.

Treat ML and APIs as operational dependencies

An ML pipeline also has running tasks, persistent effects and shared quotas. In Agent Platform, failure_policy=fast stops scheduling new tasks after failure; already scheduled tasks can continue. Do not interpret this policy as rollback of files, outputs or external calls. Define which outputs remain provisional and how a retry recognizes them. In a finished batch job, 970 successes, twenty failures and ten incomplete inputs require handling thirty inputs before claiming coverage of the original thousand. The partialFailures list can be bounded, so it does not replace counts and output reconciliation. For an API such as Vision, distinguish request quota from feature quota. In the configured example of 100 requests per minute and 600 pages per minute, with eight pages per request, the theoretical ceiling is 75 requests. These are fictional configuration values, not defaults. Planning should account for throttling, backlog, input identity and the deadline for delivering results to the business.

Deliver a change that RUN can maintain

Organize evidence by requirement and owner: post-failure capacity, update strategy, volume access, actual connectivity, connection recovery and processing completeness. For each, retain configuration, version, rehearsal, result and known limitation. If only the local calculation ran, state that rather than presenting it as a GKE test. A useful meeting update is: “Single-node maintenance fits the capacity model; zone loss does not. Acceptance remains pending additional capacity and a recovery rehearsal.” This supports time, cost and risk decisions without conflating different tests. RUN handover should identify who approves an exception, who observes backlog, who can recover connections and when a change must stop. Summarize the lesson with three questions: can the replacement run, can it serve the load, and is there evidence of that behavior under the failure we promise to tolerate? A negative answer should produce a concrete action before acceptance.

"""Original finite capacity model. No scheduler, Kubernetes API or cloud calls.
All nodes supply residual CPU/memory available for identical workload Pods.
Only integer resource fit, explicit eligibility and node removal are modeled.
"""
import json


def integer(value, name, minimum=0):
 if type(value) is not int or value < minimum:
 raise ValueError(name + ' must be an integer >= ' + str(minimum))
 return value


def capacity(nodes, cpu_m, memory_mib, unavailable=):
 integer(cpu_m, 'pod CPU', 1)
 integer(memory_mib, 'pod memory', 1)
 if not isinstance(nodes, list):
 raise ValueError('nodes must be a list')
 if not isinstance(unavailable, (list, tuple, set)):
 raise ValueError('unavailable must be a collection of IDs')
 if any(not isinstance(x, str) or not x for x in unavailable):
 raise ValueError('unavailable IDs must be nonempty strings')
 blocked = set(unavailable)
 known, slots = set, {}
 for n in nodes:
 if not isinstance(n, dict) or set(n)!= {'id', 'zone', 'cpu_m', 'memory_mib', 'eligible'}:
 raise ValueError('exact node schema required')
 if not isinstance(n['id'], str) or not n['id'] or n['id'] in known:
 raise ValueError('unique nonempty node IDs required')
 if not isinstance(n['zone'], str) or not n['zone']:
 raise ValueError('nonempty zone required')
 if type(n['eligible']) is not bool:
 raise ValueError('eligibility must be explicit boolean')
 integer(n['cpu_m'], 'node CPU')
 integer(n['memory_mib'], 'node memory')
 known.add(n['id'])
 slots[n['id']] = 0 if n['id'] in blocked or not n['eligible'] else min(n['cpu_m'] // cpu_m, n['memory_mib'] // memory_mib)
 if blocked - known:
 raise ValueError('unknown unavailable node ID')
 by_zone = {}
 for n in nodes:
 by_zone[n['zone']] = by_zone.get(n['zone'], 0) + slots[n['id']]
 return {'totalSlots': sum(slots.values), 'slotsByNode': dict(sorted(slots.items)),
 'slotsByZone': dict(sorted(by_zone.items))}


def removal_plan(nodes, replicas, cpu_m, memory_mib, unavailable=):
 integer(replicas, 'replicas')
 result = capacity(nodes, cpu_m, memory_mib, unavailable)
 return result | {'replicas': replicas, 'fitsModel': result['totalSlots'] >= replicas,
 'shortfall': max(0, replicas - result['totalSlots'])}


def voluntary_allowance(healthy, minimum_healthy):
 # Simplified arithmetic, not a reproduction of the Kubernetes PDB controller.
 integer(healthy, 'healthy')
 integer(minimum_healthy, 'minimum healthy')
 return max(0, healthy - minimum_healthy)


def node(key, zone='a', cpu_m=4000, memory_mib=8192, eligible=True):
 return dict(id=key, zone=zone, cpu_m=cpu_m, memory_mib=memory_mib, eligible=eligible)


def run:
 checks = []
 def check(name, value):
 if not value: raise AssertionError(name)
 checks.append(name)
 def reject(name, fn):
 try: fn
 except ValueError: checks.append(name)
 else: raise AssertionError(name)
 nodes = [node('a1'), node('a2'), node('b1', 'b'), node('b2', 'b')]
 baseline = removal_plan(nodes, 6, 2000, 3072)
 single = removal_plan(nodes, 6, 2000, 3072, ['a1'])
 zone = removal_plan(nodes, 6, 2000, 3072, ['a1', 'a2'])
 check('baseline eight slots', baseline['totalSlots'] == 8 and baseline['fitsModel'])
 check('one node removed six slots', single['totalSlots'] == 6 and single['fitsModel'])
 check('one zone removed four slots', zone['totalSlots'] == 4 and not zone['fitsModel'])
 check('zone loss shortfall two', zone['shortfall'] == 2)
 check('memory limits slots', capacity([node('m', cpu_m=6000, memory_mib=6144)], 2000, 4096)['totalSlots'] == 1)
 fragmented = [node('f1', cpu_m=1500), node('f2', cpu_m=1500)]
 check('aggregate CPU cannot fit a single pod', capacity(fragmented, 2000, 1024)['totalSlots'] == 0)
 check('ineligible node contributes zero', capacity([node('x', eligible=False)], 1000, 1024)['totalSlots'] == 0)
 check('zero residual CPU contributes zero', capacity([node('z', cpu_m=0)], 1000, 1024)['totalSlots'] == 0)
 check('zero residual memory contributes zero', capacity([node('z', memory_mib=0)], 1000, 1024)['totalSlots'] == 0)
 check('per-zone counts retained', baseline['slotsByZone'] == {'a': 4, 'b': 4})
 check('input ordering does not matter', capacity(list(reversed(nodes)), 2000, 3072) == capacity(nodes, 2000, 3072))
 check('empty nodes zero slots', capacity([], 1000, 1024)['totalSlots'] == 0)
 check('zero replicas fit empty model', removal_plan([], 0, 1000, 1024)['fitsModel'])
 check('all nodes removed no slots', capacity(nodes, 2000, 3072, [n['id'] for n in nodes])['totalSlots'] == 0)
 check('repeated removal ID counted once', capacity(nodes, 2000, 3072, ['a1', 'a1']) == capacity(nodes, 2000, 3072, ['a1']))
 check('one voluntary disruption in simplified model', voluntary_allowance(6, 5) == 1)
 check('no allowance below minimum', voluntary_allowance(4, 5) == 0)
 check('no allowance at minimum', voluntary_allowance(5, 5) == 0)
 reject('zero CPU request rejected', lambda: capacity(nodes, 0, 1024))
 reject('zero memory request rejected', lambda: capacity(nodes, 1000, 0))
 reject('boolean request rejected', lambda: capacity(nodes, True, 1024))
 reject('fractional request rejected', lambda: capacity(nodes, 1.5, 1024))
 reject('negative residual CPU rejected', lambda: capacity([node('n', cpu_m=-1)], 1000, 1024))
 reject('negative residual memory rejected', lambda: capacity([node('n', memory_mib=-1)], 1000, 1024))
 reject('duplicate node ID rejected', lambda: capacity([node('x'), node('x')], 1000, 1024))
 reject('unknown excluded node rejected', lambda: capacity(nodes, 1000, 1024, ['missing']))
 reject('string exclusion collection rejected', lambda: capacity(nodes, 1000, 1024, 'a1'))
 reject('nonboolean eligibility rejected', lambda: capacity([node('x', eligible=1)], 1000, 1024))
 reject('empty zone rejected', lambda: capacity([node('x', zone='')], 1000, 1024))
 reject('negative replica count rejected', lambda: removal_plan(nodes, -1, 1000, 1024))
 reject('negative healthy count rejected', lambda: voluntary_allowance(-1, 0))
 reject('boolean minimum rejected', lambda: voluntary_allowance(1, False))
 check('input nodes unchanged', nodes == [node('a1'), node('a2'), node('b1', 'b'), node('b2', 'b')])
 return {'passed': len(checks), 'checks': checks, 'baseline': baseline, 'singleNode': single,
 'zoneLoss': zone, 'network': False, 'vendorExecution': False, 'persistentWrites': False}

if __name__ == '__main__':
 print(json.dumps(run, indent=2))
IN PRACTICE

Four nodes across two zones support six Pods after one node is lost, but only four Pods after an entire zone is lost.

Common pitfalls

Treating PDB as failure prevention, pooling CPU across nodes, using RWO as Pod exclusivity, trusting DNS because PSC is accepted and treating fail fast as rollback.

Related topics: Data migration and reconciliation · Hybrid networking and storage recovery · Acceptance criteria and RUN autonomy

Take this idea with you

Calculate residual capacity, confirm eligibility and dependencies, and accept the change using evidence for the required failure scenario.

Create account

Reference: Kubernetes disruptions · Current linked standard guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.