← Professional Cloud DevOps Engineer: delivery and reliability
21 / 25 · 135 MIN

Platform contingency: access, configuration and identity

Rehearse emergency access, coordinate reconciliation and review shared identities before activating contingency environments.

1. Rehearse access under the intended failure

Contingency starts with an operational question: who can perform the first action when the normal path fails? An emergency account stored in a vault does not solve the problem if opening that vault requires unavailable SSO. In the fictional APS exercise, the incident affects identity federation and the runbook points to a portal depending on it. The plan contains a credential but has not demonstrated a usable path. Record the entire chain: operator, authentication mechanism, approval, credential retrieval, network and destination service. Define the purpose of access. Recovering federation, repairing the pipeline and performing an operational action can require different capabilities. The breakglass label does not automatically assign every role. For each purpose, identify scope, owner, use conditions and closure evidence. Exceptional access should be traceable and return to the normal model after intervention. Exact rules depend on the organization; this exercise does not describe internal BNP Paribas procedures. A useful drill deliberately removes the dependency it is intended to tolerate, in an authorized environment. Confirming access on a day when SSO works does not answer the earlier question. Retain the outcome, limitations and outstanding actions. If approval also depends on the unavailable system, that dependency needs a previously agreed operational solution, rather than improvisation during the incident.

2. Confirm identity and destination in each tool

A terminal session can contain several authentication configurations. Credentials used by gcloud are distinct from the Application Default Credentials configuration a Python library may consult. ADC first checks configuration named by GOOGLE_APPLICATION_CREDENTIALS, then the conventional local file and finally the attached account obtained through the metadata server. This order describes credential resolution, not the relative quality or security of each option. An environment variable inherited by the runner can explain why the application does not use the attached account the operator expected. Before a contingency action, confirm the process’s effective identity and selected destination. Configured project does not replace that check. Likewise, fetching GKE credentials can change kubectl’s current context. If you moved from the drill cluster to production to inspect information, a subsequent operation without an explicit context can target production. Use a check compatible with your process and retain the destination in the change record. Do not confuse obtaining cluster access information with receiving authorization to change workloads either. Permission to read metadata and generate local configuration does not establish that creation is allowed. In review, separate these questions: who presents the request, to which cluster, for what action, and which controls can still reject it? This separation helps diagnose failures without widening permissions through trial and error.

3. Coordinate manual changes with the managed source

Config Sync distinguishes webhook-based drift prevention from reconciliation that restores source-defined state. Disabling the former does not remove the latter. In a drill, the team changes a field manually, observes immediate success and declares the repair complete. Minutes later the field returns to its source value. That sequence is consistent with reconciliation; it does not prove the initial request failed or another operator contradicted the team. Identify the responsible controller and source before repeating the change. The first option should be to align the repair with the applicable managed process. If an approved exception requires temporary suspension, the runbook needs to identify the installed management method, scope to suspend, objects or configuration to retain and how to resume. Stopping an isolated Pod may be insufficient because another component manages it. Documentation distinguishes management paths for the root source and namespace-scoped sources; do not copy commands between paths without confirming actual configuration. At closure, compare approved intent, observed state and the source that will resume reconciliation. A temporary change may need to be incorporated or reversed under control. Define who makes that decision and when. The operational question is whether the next management cycle will retain intended state, rather than merely whether the last command returned success.

4. Design identity boundaries across clusters

In a fleet, some features use sameness to treat certain resources with matching names as equivalent. Consequences depend on the feature and its selectors. This lesson focuses on one concrete workload-principal form: pool, namespace and Kubernetes ServiceAccount name. If these elements match within a shared scope, changing the cluster name does not demonstrate separation of that identity. Different projects can also share scope when using fleet Workload Identity Federation as described in the scenario. To review the boundary, start with the actual grant and identify what it selects. Do not conclude that every fleet workload received every permission. Do not conclude that equal names in different pools are the same principal either. The full identifier and authorization mechanism matter. In the exercise, production and trial are trust classes supplied by the inventory author; they are neither cryptographic properties nor Google Cloud fields that automatically enforce isolation. Creation of namespaces and ServiceAccounts deserves review because it can permit creating names already selected by a grant. In a fictional scenario, a less-trusted team creates the ledger/exporter pair in a cluster using the same pool. Naming the cluster trial does not resolve the problem. Examine who controls those objects, which workloads can use them and whether grant scope matches intent. Keep data and hypothesis separate until effective access is confirmed.

5. Include identities outside the membership list

A membership list answers the administrative question of fleet registration. It may not answer the security question of every workload sharing a pool. Documentation describes clusters in the host project, including a non-member, with Workload Identity Federation using the same pool. Identities matching the namespace and ServiceAccount selector can matter even outside the initial list. This possibility makes inventory scope part of the review rather than a simple export step. In the local worksheet, each record has an identifier, cluster, project, pool, namespace, ServiceAccount, trust class and membership flag. Project and membership remain context but do not separate the triple selected for comparison. A record without a confirmed pool uses null. The program marks inventory as incomplete; it neither invents a separate pool nor inherits the preceding row’s value. The same care applies to an unknown trust class. When receiving inventory from several teams, confirm data origin and observation date before using the result in a change. The exercise queries no services to check freshness. A report without overlaps describes only supplied rows and the implemented rule. For the committee, present scope, gaps, findings and who will confirm missing points. Do not turn missing information into a guarantee of isolation.

6. Recover configuration compatible with the target

An old configuration revision can be identifiable yet incompatible with the destination. A package containing a Deployment in extensions/v1beta1 cannot depend on a server later than Kubernetes 1.16 that no longer serves that API. Migration to a supported API requires reviewing fields and behavior, including required selectors. Editing a version label alone is not enough to assume equivalence. Use the intended destination during the drill and check object interpretation before connecting the application to real traffic. CustomResourceDefinitions add another dependency. The definition establishes the API for custom resources; stored objects may contain state needed for recovery. Deleting a CRD also deletes associated objects, and recreating it does not automatically restore them. A proposal to “clean up and reinstall” therefore needs analysis of what will be removed, the recoverable copy and the validation path. An identically named new object does not establish continuity of previous data. Connect these dependencies to the bootstrap plan. If the pipeline creating the network runs only on runners inside that nonexistent network, retaining code and state is insufficient. An authorized environment capable of starting the work is needed. In the exercise, the team first designs that path and then rehearses manifest compatibility. Record the actual sequence, with owners for identity, networking and configuration, so another person can execute the plan.

7. Treat admission as an operational dependency

An authenticated and authorized identity can submit a request that later fails admission. In the scenario, a webhook matches the request, no exclusions apply and the call times out. With failurePolicy: Fail, the call error prevents accepting that request. Repeating authentication or granting more roles does not repair the unavailable endpoint. Investigation should distinguish explicit policy rejection, communication error and failure at an earlier control. This distinction matters during recovery because the admission service also needs an operational path. If the application being restored is required by the webhook admitting that very restore, the plan can contain another circular dependency. Before the window, review scope, availability and the authorized procedure for that failure mode. Do not generically change policy to Ignore merely to obtain a green result: that can remove a control the organization requires. The solution must preserve the control’s objective and be rehearsed within approved scope. In the incident record, retain a concrete observation: requested operation, destination, rejection stage and received error. Reporting “missing permissions” when an admission timeout occurred directs the team toward the wrong repair. For shift handover, identify the actual blocker, dependency owner and next validated action. Diagnosis should reduce uncertainty before expanding privileges or changing configuration.

8. Run the overlap inventory

Run python3 run.py using Python 3.13 or compatible. The code uses only the standard library and prints JSON to the terminal. It does not contact Google Cloud, read IAM policies or change resources. It compares the exact pool, namespace and ServiceAccount triple in supplied records. Scope assumes the principal form based on these names; it implements neither UID-based subjects, every IAM selector type nor condition evaluation. Pool aliases represent complete unique identifiers including the host project; short pool names in different projects are insufficient for this comparison. These are fictional data, not real credentials. Start with the nine named cases. Mixed-trust produces a group containing different known classes. Same-trust-sharing shows a shared identity without a declared class difference. Nonmember-overlap retains the non-member in analysis. Unknown-pool and unknown-trust leave the conclusion incomplete. Then compare different-pool, different-namespace and different-service-account to understand which dimensions define the key. No result sets accessProven or productionAuthorized to true. The program explores 5,184 pairs of combinations and rejects 14 invalid inputs. It also confirms that reversing pairs and testing all six orders of three records preserves the result and input data is not modified. These checks demonstrate properties of the comparator, not completeness of actual inventory. As a task, add a third cluster, justify its trust class and explain what evidence you would request before approving the change. An overlap is an investigation point; absence of findings is not authorization.

"""Original inventory exercise for name-based workload principals; not an IAM evaluator."""
from copy import deepcopy
from itertools import product, permutations
import hashlib
import json
from pathlib import Path

FIELDS = {'id', 'cluster', 'project', 'pool', 'namespace', 'serviceAccount', 'trust', 'fleetMember'}


def inspect_inventory(records):
 """Compare supplied pool/namespace/serviceAccount tuples exactly.

 Scope assumes the namespace/name principal form, not UID-based subjects or
 all IAM principal selectors. Pool and trust may be unknown (None). Project
 and fleet membership remain inventory context, not separators for a shared
 pool. Findings require review; no actual access or effective policy is read.
 """
 if not isinstance(records, list) or not records:
 raise ValueError('nonempty inventory required')
 ids, groups, unknown = set, {}, []
 for record in records:
 if not isinstance(record, dict) or set(record)!= FIELDS:
 raise ValueError('exact inventory fields required')
 for field in FIELDS - {'fleetMember'}:
 value = record[field]
 if value is None and field in {'pool', 'trust'}:
 continue
 if not isinstance(value, str) or not value.strip:
 raise ValueError('nonempty string required: ' + field)
 if type(record['fleetMember']) is not bool:
 raise ValueError('membership must be boolean')
 if record['id'] in ids:
 raise ValueError('duplicate record ID')
 ids.add(record['id'])
 missing = [field for field in ('pool', 'trust') if record[field] is None]
 if missing:
 unknown.append({'id': record['id'], 'fields': missing})
 if record['pool'] is not None:
 key = (record['pool'], record['namespace'], record['serviceAccount'])
 groups.setdefault(key, []).append(record)
 shared = []
 for key, members in sorted(groups.items):
 if len(members) < 2:
 continue
 trusts = sorted({r['trust'] for r in members if r['trust'] is not None})
 shared.append({'principal': dict(zip(('pool', 'namespace', 'serviceAccount'), key)),
 'ids': sorted(r['id'] for r in members),
 'clusters': sorted({r['cluster'] for r in members}),
 'projects': sorted({r['project'] for r in members}),
 'knownTrusts': trusts,
 'mixedKnownTrust': len(trusts) > 1,
 'unknownTrust': any(r['trust'] is None for r in members),
 'includesNonMember': any(not r['fleetMember'] for r in members)})
 return {'records': len(records), 'sharedPrincipals': shared,
 'mixedTrustGroups': sum(g['mixedKnownTrust'] for g in shared),
 'unknownRecords': sorted(unknown, key=lambda r: r['id']),
 'inventoryComplete': not unknown,
 'accessProven': False, 'productionAuthorized': False}


def workload(name, **changes):
 return {'id': name, 'cluster': 'cluster-' + name, 'project': 'project-' + name,
 'pool': 'pool-a', 'namespace': 'funds', 'serviceAccount': 'processor',
 'trust': 'production', 'fleetMember': True, **changes}


def evidence:
 a = workload('a')
 variants = [
 ('mixed-trust', workload('b', trust='trial')),
 ('different-pool', workload('b', pool='pool-b', trust='trial')),
 ('different-namespace', workload('b', namespace='trial', trust='trial')),
 ('different-service-account', workload('b', serviceAccount='trial', trust='trial')),
 ('same-trust-sharing', workload('b')),
 ('nonmember-overlap', workload('b', trust='trial', fleetMember=False)),
 ('unknown-pool', workload('b', pool=None, trust='trial')),
 ('unknown-trust', workload('b', trust=None)),
 ]
 fixtures = []
 for name, b in variants:
 inventory = [a, b]
 before = deepcopy(inventory)
 result = inspect_inventory(inventory)
 assert inventory == before
 assert inspect_inventory(list(reversed(inventory))) == result
 fixtures.append({'id': name, **result})
 assert fixtures[0]['mixedTrustGroups'] == 1
 assert all(f['mixedTrustGroups'] == 0 for f in fixtures[1:5])
 assert len(fixtures[4]['sharedPrincipals']) == 1
 assert fixtures[5]['sharedPrincipals'][0]['includesNonMember']
 assert fixtures[5]['mixedTrustGroups'] == 1
 assert not fixtures[6]['inventoryComplete'] and not fixtures[7]['inventoryComplete']
 triple = [a, workload('b', trust='trial'), workload('c', trust=None, fleetMember=False)]
 triple_before = deepcopy(triple)
 triple_result = inspect_inventory(triple)
 assert triple_result['mixedTrustGroups'] == 1
 assert triple_result['sharedPrincipals'][0]['unknownTrust']
 assert not triple_result['inventoryComplete']
 for ordering in permutations(triple):
 assert inspect_inventory(list(ordering)) == triple_result
 assert triple == triple_before
 fixtures.append({'id': 'mixed-and-unknown', **triple_result})
 values = list(product(('pool-a', 'pool-b', None), ('funds', 'trial'),
 ('processor', 'observer'), ('production', 'trial', None),
 (True, False)))
 combinations = 0
 for left, right in product(values, repeat=2):
 def row(name, values):
 return workload(name, **dict(zip(('pool','namespace','serviceAccount','trust','fleetMember'), values)))
 result = inspect_inventory([row('a', left), row('b', right)])
 same = left[0] is not None and right[0] is not None and left[:3] == right[:3]
 mixed = same and left[3] is not None and right[3] is not None and left[3]!= right[3]
 assert len(result['sharedPrincipals']) == int(same)
 assert result['mixedTrustGroups'] == int(mixed)
 assert result['inventoryComplete'] == all(x[0] is not None and x[3] is not None for x in (left,right))
 assert not result['accessProven'] and not result['productionAuthorized']
 combinations += 1
 invalid = [[], {}, [a, a], [{**a, 'id': ''}], [{**a, 'cluster': None}],
 [{**a, 'pool': ''}], [{**a, 'trust': ' '}], [{**a, 'fleetMember': 1}],
 [{**a, 'namespace': None}], [{**a, 'serviceAccount': []}],
 [{k:v for k,v in a.items if k!= 'project'}], [{**a, 'role': 'admin'}],
 [{**a, 'project': False}], [None]]
 for records in invalid:
 try:
 inspect_inventory(records)
 except ValueError:
 pass
 else:
 raise AssertionError('invalid input accepted')
 return {'fixtures': fixtures, 'pairCombinations': combinations, 'invalidInputs': len(invalid),
 'inputPreserved': True, 'orderIndependent': True, 'triplePermutations': 6, 'network': False,
 'cloudExecuted': False, 'iamPoliciesRead': False, 'persistentWrites': False,
 'independentVerification': False,
 'scriptSha256': hashlib.sha256(Path(__file__).read_bytes).hexdigest}


if __name__ == '__main__':
 print(json.dumps(evidence, sort_keys=True, indent=2))
IN PRACTICE

Two differently named clusters can share the same principal selected by pool, namespace and ServiceAccount; membership does not necessarily bound the entire relevant inventory.

Common pitfalls

SSO in the emergency path; assuming ADC identity from CLI; stale kubectl context; disabled webhook treated as suspended reconciliation; cluster names as an IAM boundary.

Related topics: Recovery evidence and dependencies · Environment management and change control · Pipeline identity and security

Take this idea with you

Contingency requires usable access, a confirmed destination, recoverable configuration and identities with understood and checked scope.

Create account

Reference: Operations best practices · Current linked guide; edition date unconfirmed (2026-09-30 inspection)

Google Cloud is a trademark of Google LLC. bigsavant.com is an independent preparation platform and is not affiliated with, associated with, sponsored, authorised or endorsed by Google. Content and questions are original, are not official exam questions, and completing our tests does not award or guarantee any certification. Names are used only to identify the subject. All other trademarks belong to their respective owners.