Rehearse the complete recovery chain
At 07:40, an APS team at a fictional bank prepares to open its funds service. The IdP is unavailable and the on-call administrator has lost their second factor. The manager must establish whether recovery access works under the incident conditions. This case does not describe internal BNP Paribas procedures. Draw a chain containing the operator, identity, factor, device, network path, vault and authorization. Mark every dependency on the failed component. An independent account is insufficient if its only spare key is inside a vault requiring the unavailable IdP. Administrative recovery also has privilege boundaries. Google distinguishes recovery of ordinary users from recovery of another administrator account, which requires another super administrator. Confirm the primary address in the runbook: receiving messages through an alias does not make it the Google sign-in identifier. Keep recovery compatible with applicable 2SV policy; moving the account into a unit without enforcement changes its protection. During rehearsal, the observer records where the operator became blocked, which evidence they consulted and who can decide the next step. Include operator absence and an expired session in the exercise. Do not place real recovery codes in the shared report. As a guided decision, compare two plans: adding another operator dependent on the same vault, or providing an authorized independent recovery path. Explain which dependency each plan removes and which dependency remains.
Restore federation without reopening exposure
A monthly reconciliation process depends on a workload identity pool that appears inactive during the day. Before retirement, identify the actual consumer schedule. Reversible disabling, with an observation window covering the monthly process, helps expose dependencies. No alerts during one afternoon does not cover that process. Record the reversal condition, batch owner and expected evidence of a new token exchange. Then distinguish administrative operations. Disabling a provider prevents new exchanges, but existing tokens can continue granting access. Deleting the pool behaves differently: earlier credentials are not revoked, although they do not grant access while the pool is deleted. If the pool is undeleted before those credentials expire, they can grant access again. Restoration therefore needs an exposure assessment as well as an availability assessment. A deleted pool name remains unavailable for reuse until deletion becomes permanent. In the guided case, continuity staff want immediate restoration to meet the batch deadline. Security is still investigating a credential issued before deletion. Prepare a decision note with three entries: business dependency, restoration’s effect on that credential, and the condition still needing confirmation. Do not conclude that a credential expired merely because the console stopped displaying the pool. A new authentication test also does not establish the state of all earlier credentials. Approval should identify the consumers and residual risk that were actually assessed.
Coordinate rotation between signer and verifier
The authentication path has two sides that can change at different times. In a SAML profile with two verification certificates, overlap allows preparation of new trust before the IdP starts signing with the new key. The inspected Google guidance calls for uploading the second certificate and waiting 24 hours for account updates before switching the IdP. Plan the window around that dependency rather than treating upload as immediate effect. Retire earlier material after confirming legitimate operation and the scope of affected profiles. For OIDC using manually uploaded JWKS, the configured copy likewise requires coordinated updates. Publishing a new key at the IdP does not update that copy by itself. With public discovery, the JWKS endpoint has TLS requirements; an exception accepted in an operator’s browser does not make a self-signed certificate supported by the Google consumer. Separate the key validating a token signature from the certificate authenticating the HTTPS endpoint. Build an original transition table: signer state, keys accepted at the destination, positive test and negative test. The intermediate state should accept the intended legitimate flow while retaining rejection of untrusted material. During an incident, first confirm which profile and key set are effective. Avoid extending the change to every profile because one failed. At closure, retain configuration references and results without including private keys or real tokens.
Validate fallback destination and trust
The team receives a federation configuration file in an incident ticket. It contains no private key, but directs the client to authentication endpoints. Review its contents against the approved configuration, including audience and network destinations. Valid JSON syntax and a familiar provider name do not establish file trust. Treat configuration provenance and integrity as part of recovery. A proxy terminates TLS in front of the IdP. After changing the backend, its health check succeeds and the frontend certificate remains valid. The team still needs to establish that the backend serves the intended issuer’s keys. A document with the expected fields can contain different material. Similarly, a valid signature does not resolve a changed audience: confirm the token’s intended recipient before widening accepted audiences. Now consider an approved regional requirement for token exchange. The inspected documentation identifies regional STS endpoints as Preview. Switching token_url to the global endpoint might help availability, but requires evaluating the requirement and authorization for that deviation. Here the requirement is a fictional organizational assumption, not a legal obligation that this course assigns to every bank. Prepare two options for the decision owner: retain the restriction and defer processing, or formally assess a scoped fallback. Explain impact and missing evidence without promising that the same configuration suits every context.
Reverse a change while preserving concurrent work
At 10:00, snapshot A contains a conditional batch binding. At 10:05, a mistaken change removes the condition. At 10:08, another team adds a legitimate recovery grant. Restoring all of A would erase that grant. The requested scope is reversal of the mistaken change while preserving subsequent authorized work. Compare the earlier state, the change result and current state. If the same binding changed again, flag a conflict for reconciliation. In a real IAM operation, reading must retain conditions and writing must use concurrency protection. An older view containing a _withcond_ suffix does not authorize creation of a parallel unconditional binding. Obtain a version 3 representation and preserve relevant content. The etag accompanies the observed revision; after ABORTED, repeat reading, analysis and writing. Pasting a recent etag into an old backup does not reconcile that backup. The change manager prepares a comparison showing access to recover, the concurrent change to preserve and validation criteria. Ask another participant to explain the expected effect before execution. A version 1 response after authorized removal of the last condition can match the intended content; inspect content rather than classifying the number alone as failure. The following exercise practices this reasoning with synthetic records. Its result is a review proposal rather than an IAM payload ready for application.
Connect role state with its grants
A team disabled a custom role during containment. Bindings remain visible and the handover report says the action failed. That conclusion confuses association existence with permission effect. In the DISABLED state, bindings of that role do not grant its permissions. This does not establish the absence of other access paths. The inventory also needs to identify other roles and identities relevant to the operation. If the role was deleted two days earlier, undelete recovery can be assessed within the documented seven-day window. Earlier bindings remain recorded without effect while the role is deleted, and can regain effect on restoration. Before reactivation, review members, scope and definition permissions. The same display title on a new role does not establish continuity with the earlier identifier. Apply the reasoning to a fictional position-closing incident. The manager receives a request to restore the service’s role quickly. Produce a short record containing the containment reason, definition to recover, dependent consumers and agreed tests. Include an access check that should work again and one that should remain denied. If evidence about a binding is missing, assign a review owner instead of assuming that a working application proves correctness of all access. Record separately who approved restoration and who observed the technical result.
Separate testing, inheritance and observed effect
An organization policy proposal is in dry-run. A log shows liveResult=ALLOWED and dryRunResult=DENIED. The observed request was allowed by the live policy; the other field describes the simulated evaluation. A dashboard counting every DENIED as a block mixes outcomes with different meanings. Use the rehearsal to identify operations a future change would affect, with a business owner and remediation plan. When authorization covers only testing, limit the update to dryRunSpec and confirm that spec was preserved. A file named test does not constrain an API write. Keep the comparison between earlier and later configuration and connect results to the version tested. Configuration promotion is an additional decision requiring approved scope. To recover a blocked release, the team might propose deleting the project’s local policy. If the folder has an explicit policy, the project inherits that configuration; deletion does not necessarily restore the constraint default. Confirm current hierarchy and effective policy. After an accepted correction, consider propagation before concluding that every restriction must be removed. The guided case requires a three-sentence note: which operation fails, which policy should govern it, and which observation will demonstrate recovery. Avoid using only a successful change response as proof that behavior already changed.
Exercise: compare before, after and now
Run the supplied Python program locally with fictional data. Before executing it, predict preserve-concurrent-grant: the earlier condition should return and the recovery grant should remain. In same-binding-concurrent-edit, the current binding matches neither the earlier state nor the change result; the proposal should be absent and conflict explicit. In atomic-conflict, one conflict prevents the complete proposal even if another binding could be reversed separately. The localId field exists only in this exercise and is not an IAM field. Conditions are opaque text without CEL interpretation; logically equivalent expressions can produce conflict. Member ordering does not change comparison, but duplicates and unknown fields are rejected. The schema excludes auditConfigs and other parts of a real policy that a production tool would need to preserve. Missing bindings mean a declared empty list rather than an incomplete read silently accepted. The current etag accompanies the proposal, but the program does not verify its freshness against a server. It does not calculate effective access, contact cloud services or authorize a write. Change a condition in current state, retain an independent grant and explain the output before repeating tests. As the final deliverable, write a handover note stating the change to reverse, concurrent work preserved, unresolved conflict and next real verification needed. The hash identifies executed code; it does not authenticate snapshots or prove operational approval.
"""Offline worksheet: reconcile synthetic binding records, never call IAM.
localId is an exercise identifier, NOT an IAM binding field. Conditions are
opaque text: this program does not parse CEL or calculate effective access.
A proposal needs human review and a fresh real read-modify-write cycle.
"""
from copy import deepcopy
from hashlib import sha256
from itertools import permutations, product
from pathlib import Path
import json
def nonblank(value):
return isinstance(value, str) and bool(value.strip) and value == value.strip
def validate(snapshot):
if not isinstance(snapshot, dict) or set(snapshot)!= {'resource', 'etag', 'version', 'bindings'}:
raise ValueError('Expected exactly resource, etag, version and bindings')
if not all(nonblank(snapshot[k]) for k in ['resource', 'etag']):
raise ValueError('Resource and declared etag must be nonblank strings')
if type(snapshot['version']) is not int or snapshot['version'] not in (1, 3):
raise ValueError('Worksheet accepts schema version 1 or 3 only')
if not isinstance(snapshot['bindings'], list):
raise ValueError('Bindings must be a list')
indexed = {}
for binding in snapshot['bindings']:
if not isinstance(binding, dict) or set(binding)!= {'localId', 'role', 'members', 'condition'}:
raise ValueError('Expected exact synthetic binding fields')
if not all(nonblank(binding[k]) for k in ['localId', 'role']):
raise ValueError('Local ID and role must be nonblank strings')
if binding['localId'] in indexed:
raise ValueError('Duplicate local ID')
members = binding['members']
if not isinstance(members, list) or not members or not all(nonblank(m) for m in members):
raise ValueError('Members must be a nonempty list of nonblank strings')
if len(members)!= len(set(members)):
raise ValueError('Duplicate member')
condition = binding['condition']
if condition is not None and (not nonblank(condition) or snapshot['version']!= 3):
raise ValueError('Opaque condition text requires schema version 3')
indexed[binding['localId']] = {**deepcopy(binding), 'members': sorted(members)}
return indexed
def reconcile(before, after, current):
"""Produce an all-or-nothing structural proposal from declared snapshots.
IDs touched between before/after define the scope. A touched current record
must match after (invert it) or before (already reverted); otherwise conflict.
Unrelated current records survive. Etag is carried as evidence, not verified.
"""
old, changed, now = [validate(s) for s in (before, after, current)]
if len({s['resource'] for s in (before, after, current)})!= 1:
raise ValueError('Snapshots must describe the same declared resource')
touched = sorted(k for k in old.keys | changed.keys if old.get(k)!= changed.get(k))
proposal = deepcopy(now)
decisions = []
for key in touched:
if now.get(key) == old.get(key):
action = 'already-reverted'
elif now.get(key) == changed.get(key):
action = 'restore-before'
if key in old:
proposal[key] = deepcopy(old[key])
else:
proposal.pop(key, None)
else:
action = 'conflict'
decisions.append({'localId': key, 'action': action})
conflict = any(d['action'] == 'conflict' for d in decisions)
result = {
'resource': current['resource'], 'observedEtag': current['etag'],
'decisions': decisions, 'conflict': conflict,
'untouchedCurrentIds': sorted(set(now) - set(touched)),
'proposal': None, 'humanReviewRequired': True,
'etagFreshnessVerified': False, 'celEvaluated': False,
'effectiveAccessVerified': False, 'productionWriteAuthorized': False,
}
if not conflict:
bindings = [proposal[k] for k in sorted(proposal)]
result['proposal'] = {'resource': current['resource'], 'etag': current['etag'],
'version': 3 if any(b['condition'] is not None for b in bindings) else current['version'],
'bindings': bindings}
return result
def binding(key='batch', member='serviceAccount:batch@example.invalid', condition='window-A'):
return {'localId': key, 'role': 'roles/example.reader', 'members': [member], 'condition': condition}
def snapshot(bindings=None, etag='current-C', version=3):
return {'resource': 'projects/fictional-funds', 'etag': etag, 'version': version,
'bindings': deepcopy([binding] if bindings is None else bindings)}
def evidence:
before = snapshot(etag='before-A')
after = snapshot([binding(condition=None)], 'after-B')
concurrent = binding('recovery', 'group:recovery@example.invalid', None)
fixtures = []
def record(label, b, a, c):
original = deepcopy((b, a, c))
result = reconcile(b, a, c)
assert (b, a, c) == original
fixtures.append({'id': label, **result})
return result
r = record('preserve-concurrent-grant', before, after, snapshot([binding(condition=None), concurrent]))
assert r['proposal']['bindings'] == [binding, concurrent]
assert r['proposal']['etag'] == 'current-C'
assert r['untouchedCurrentIds'] == ['recovery']
r = record('same-binding-concurrent-edit', before, after, snapshot([binding(condition='window-C')]))
assert r['conflict'] and r['proposal'] is None
r = record('already-reverted', before, after, before)
assert r['decisions'][0]['action'] == 'already-reverted'
r = record('restore-deleted-binding', before, snapshot([]), snapshot([concurrent]))
assert r['proposal']['bindings'] == [binding, concurrent]
r = record('remove-added-binding', snapshot([]), before, snapshot([binding, concurrent]))
assert r['proposal']['bindings'] == [concurrent]
r = record('deleted-by-another-change', before, after, snapshot([]))
assert r['conflict']
r = record('version-three-restored', before, after, snapshot([binding(condition=None)], version=1))
assert r['proposal']['version'] == 3
r = record('no-declared-binding-change', before, before, snapshot([concurrent]))
assert r['decisions'] == [] and r['proposal']['bindings'] == [concurrent]
multi_before = snapshot([binding, binding('second')])
multi_after = snapshot([binding(condition=None), binding('second', condition=None)])
r = record('atomic-conflict', multi_before, multi_after,
snapshot([binding(condition=None), binding('second', condition='other-change')]))
assert r['proposal'] is None
r = record('opaque-condition-change', before, after, snapshot([binding(condition='window-A && true')]))
assert r['conflict'] # No claim that logically equivalent expressions compare equal.
r = record('member-order-only', snapshot([{'localId':'x','role':'roles/example.reader','members':['user:a','user:b'],'condition':None}]),
snapshot([{'localId':'x','role':'roles/example.reader','members':['user:b','user:a'],'condition':None}]), snapshot([]))
assert r['decisions'] == []
r = record('no-condition-version-preserved', snapshot([], version=1), snapshot([concurrent]), snapshot([concurrent]))
assert r['proposal']['version'] == 3
states = [None, binding, binding(condition=None), binding(condition='window-C')]
combinations = 0
for b, a, c in product(states, repeat=3):
inputs = [snapshot([] if x is None else [x]) for x in (b, a, c)]
result = reconcile(*inputs)
expected_conflict = b!= a and c!= b and c!= a
assert result['conflict'] == expected_conflict
if not expected_conflict:
expected = c if b == a else b
assert result['proposal']['bindings'] == ([] if expected is None else [expected])
combinations += 1
rows = [binding(condition=None), concurrent, binding('observer', condition=None)]
expected = reconcile(before, after, snapshot(rows))
orders = 0
for order in permutations(rows):
assert reconcile(before, after, snapshot(list(order))) == expected
orders += 1
invalid = []
def altered(fn):
s = snapshot; fn(s); invalid.append(s)
altered(lambda s:s.pop('etag'))
altered(lambda s:s.update(etag=''))
altered(lambda s:s.update(etag=' stale '))
altered(lambda s:s.update(resource=''))
altered(lambda s:s.update(version=True))
altered(lambda s:s.update(version=2))
altered(lambda s:s.update(version=1))
altered(lambda s:s.update(bindings={}))
altered(lambda s:s['bindings'].append(deepcopy(s['bindings'][0])))
altered(lambda s:s['bindings'][0].update(localId=''))
altered(lambda s:s['bindings'][0].update(role=''))
altered(lambda s:s['bindings'][0].update(members=[]))
altered(lambda s:s['bindings'][0].update(members=['user:a','user:a']))
altered(lambda s:s['bindings'][0].update(members=['']))
altered(lambda s:s['bindings'][0].update(condition=''))
altered(lambda s:s['bindings'][0].update(condition={'expression':'true'}))
altered(lambda s:s['bindings'][0].update(extra='unsupported'))
altered(lambda s:s.update(auditConfigs=[]))
invalid.extend([None, [], 'policy'])
for value in invalid:
try:
reconcile(before, after, value)
except ValueError:
pass
else:
raise AssertionError('Invalid input accepted')
other = snapshot; other['resource'] = 'projects/other'
try:
reconcile(before, after, other)
except ValueError:
pass
else:
raise AssertionError('Resource mismatch accepted')
return {'scriptSha256': sha256(Path(__file__).read_bytes).hexdigest,
'fixtures': fixtures, 'stateCombinations': combinations, 'inputPermutations': orders,
'invalidInputs': len(invalid)+1, 'inputPreserved': True, 'orderIndependent': True,
'network': False, 'cloudExecuted': False, 'persistentWrites': False}
if __name__ == '__main__':
print(json.dumps(evidence, ensure_ascii=False, indent=2))
Snapshot A contains a condition; B removes it; C adds a legitimate grant. The proposal restores the condition and retains the grant. If C also changed the same binding, reconciliation is required.
Common pitfalls
Confusing disabling with revocation; restoring a pool without assessing earlier tokens; copying a recent etag into stale content; treating dry-run as a block; deleting local policy while ignoring inheritance.
Related topics: Revocation, sessions and recovery decisions · Identities, credentials and access evidence · Governance, scope and control evidence
Recovery requires restoring legitimate operation, retaining trust boundaries and demonstrating that rollback respects current state.
Reference: Recover an account protected by 2-Step Verification · Current linked guide; edition date unconfirmed (2026-09-30 inspection)