1. Reconstruct delivery before intervening
At 22:10, a delivery team hands APS the message “the release failed.” That sentence does not identify what remains installed or what is still running. Start a timeline containing release, target, rollout, phase, job and job-run identities. Add timestamps and evidence references. A job may have several attempts, and an attempt may create an external effect before losing its acknowledgement. Record observed state separately from intended action. “Verify failed because the endpoint was unreachable” is an observation; “retry after fixing the route” is a decision needing an owner and an entry condition. In this lesson's exercise, prod-west and prod-east share a support team but retain different delivery histories. A release succeeding in east does not establish that west received it. Handover should let another team reconstruct events without depending on the operator's memory. Include what remains unknown: an external acknowledgement, converging traffic or a running job. Do not fill those gaps with the last green dashboard state. This discipline supports choosing whether to observe, retry, interrupt or recover without turning an assumption into a fact.
2. Choose recovery per target
A Cloud Deploy rollback creates a new rollout from an earlier release. Its default selection uses the last success on the specified target; a particular release can also be selected. In our example, r18 succeeded in prod-west, r19 succeeded only in staging and r20 failed in prod-west. West's default reference is r18. Before execution, the team confirms that the release remains eligible and compatible with the current environment. Historical success is an operational reference, not a guarantee that dependencies, data and policies are unchanged. Retain the new rollout and link it to the incident motivating recovery. Do not replace old results with new ones. The acceptance objective in this exercise is to demonstrate that the critical workflow works again on the affected target, while reconciling effects produced during failure. If an external system received instructions while r20 was active, returning to r18 does not erase them. The plan must identify who checks external state and who may authorize compensation. Rollback and data decisions may have different owners and durations, which need coordination.
3. Separate abandonment, cancellation and termination
Abandoning a release prevents new deliveries from that release and is permanent. Existing rollouts, including queued ones, are not cancelled by abandonment. Therefore, if the incident requires containment, inspect existing rollouts too. An abandoned resource cannot be reactivated as a recovery option. Update the runbook's eligible-release list when withdrawing a version; otherwise the night team may discover that its only candidate is unavailable during an incident. Cancelling a rollout can leave it CANCELLING while outstanding job runs finish. Terminating an active job run leaves its job failed. None of these states demonstrates that external effects were undone. In a predeploy creating a reservation, that reservation may survive interruption. Before another attempt, look up the result using the logical operation identity. If supported by evidence, report “cancellation requested; job X still active; reservation Y awaiting reconciliation.” That statement assigns concrete work. “Everything reverted” would prematurely close tasks that remain necessary. Keep the distinction visible in the incident timeline and in the next shift's entry checklist, especially when application state and deployment state are owned by different teams.
4. Coordinate retries and manual intervention
Automatic repair needs a time budget and an ownership rule for human intervention. In this model, three retries wait 2, 4 and 8 minutes, and each attempt takes 1 minute. That is 17 minutes until the third retry ends, excluding queues and control latency. If the window has 12 minutes left, the plan already exceeds its stated assumptions. Reducing waits just to fit the window may worsen a saturated dependency. Discuss mitigation, an authorized extension or safe interruption with the responsible people first. Manually retrying a job aborts the ongoing repairRolloutRule run. Do not let two teams assume the other is still following the original automatic sequence. Record the intervention, observed state and owner of the next decision. The disableRollbackIfRolloutPending option prevents that rollback when another rollout is pending on the target. It addresses a defined concurrency condition but does not prove the pending delivery is safe. The candidate still needs assessment and the target needs coordination. In this exercise, recovery authority belongs to the named incident owner; it is not inferred simply because an automation exists.
5. Build hooks that recognize the same intent
The hook author is responsible for idempotency. First identify the business effect that should occur once: reserving capacity for prod-west and r70, for example. A job-run ID changes between attempts and cannot, by itself, recognize that intent. The Python exercise uses a stable key, payload and receipt. If the first call records the effect but loses its ACK, repeating the same key and payload returns the existing receipt. A different key creates another reservation, illustrating why observed failure and absence of effect are not equivalent. The same identifier with a different payload is treated as a conflict. Moving from 500 to 700 units needs a new decision; it must not be disguised as a retry or silently return success with the old value. The model is sequential and stores data only in memory. It does not prove concurrency exclusion, crash durability, authentication or atomicity with an external service. A real integration needs these properties designed and exercised with its provider. Use the exercise to explain intent identity and recognize duplication; do not use it as a production operations library.
6. Check order, configuration and boundaries
If verification needs synthetic data, the data must exist before verify. Putting preparation in postdeploy creates an inverted dependency, because verify precedes postdeploy when configured. In automated canary, predeploy runs in the first phase and postdeploy in the last; they are not tasks repeated at each percentage. A task list is sequential and stops at its first failure. If reserve passed and seed failed halfway, announce did not run; reserve's effects and partial data still need reconciliation. Per-target configuration also deserves explicit review. matchTargetLabels combines labels with AND: region=west and class=critical requires both. Two different values of the same parameter, without distinguishing selectors, are not automatically distributed by target order. In the exercise, the team must show each child's rendered manifest and confirm its expected replica count. Bind that evidence to target and release. Approval of an intent table does not establish that the delivered file contains the right values. If there is a mismatch, hold normal promotion and fix selection before collecting fresh evidence.
7. Apply freezes and observe actual traffic
A deployment policy can restrict actions such as CREATE or ADVANCE during its configured window. A release may exist while creating its rollout is prohibited. Current policy is evaluated when the action is attempted; creating the release before a freeze grants no exemption. Conditions within a selector combine with AND; separate selectors combine with OR. Check the invoker and time zone too. Permission to override policy does not replace the operation's IAM permission. In this exercise, an exception needs the team's defined authorization as well as technical ability to perform it. After a Cloud Run traffic change, observe convergence and effects from in-flight requests. A late response from the previous revision during transition does not, by itself, establish command failure. A revision allocated zero in the traffic split can also remain directly accessible through a tag for authorized testing. For a trial without business effects, control its data, endpoints and identities. Writing zero on a diagram does not isolate a shared database. Recovery criteria should measure the relevant workflow and track external results, rather than only desired allocation state.
8. Practise, explain and hand over to RUN
Run the local Python model and compare lost-ack-stable-key with lost-ack-fresh-key. The first sequence retains one 500-unit reservation; the second ends with two, totaling 1000. In changed-intent, the 700-unit request under the previous key is rejected and the total remains 500. Before reading answers, predict each result and identify evidence that would change your decision. Then vary the retry count without changing intent. The expected outcome remains a single effect. The exercise also checks invalid inputs and copies payloads to avoid accidental mutation of stored records. Finish by preparing an English handover note: what was requested, what was observed, which external effect is confirmed, who owns recovery and what criterion permits closure. Attach the correct identifiers and state what still needs checking. In a tabletop exercise, another person should be able to choose the next step using only that note and its references. The final challenge is to justify why abandonment, cancellation, termination and rollback address different problems. Assessment uses fictional situations and explicit rules; it does not describe an actual bank's internal procedures or replace technical validation of a production integration.
"""Original sequential teaching model: not a Cloud Deploy or payment emulator.
No network or persistent writes. A production implementation needs durable,
atomic idempotency records, authorization and reconciliation with its provider.
"""
from copy import deepcopy
from hashlib import sha256
import json
from pathlib import Path
def reserve(ledger, key, payload, *, lose_ack=False):
if not isinstance(key, str) or not key.strip:
raise ValueError('nonempty operation key required')
if not isinstance(payload, dict) or set(payload)!= {'target', 'units'}:
raise ValueError('target and units required')
if not isinstance(payload['target'], str) or not payload['target'].strip:
raise ValueError('nonempty target required')
if type(payload['units']) is not int or payload['units'] <= 0:
raise ValueError('positive integer units required')
if type(lose_ack) is not bool:
raise ValueError('boolean acknowledgement flag required')
if key in ledger:
if ledger[key]['payload']!= payload:
return {'status': 'conflict', 'created': False, 'receipt': None}
return {'status': 'replayed', 'created': False,
'receipt': ledger[key]['receipt']}
receipt = f"reservation-{len(ledger) + 1}"
ledger[key] = {'payload': deepcopy(payload), 'receipt': receipt}
return {'status': 'unknown' if lose_ack else 'created', 'created': True,
'receipt': None if lose_ack else receipt}
def totals(ledger):
return sum(row['payload']['units'] for row in ledger.values)
def fixture(name, calls):
ledger = {}
results = [reserve(ledger, key, payload, lose_ack=lost)
for key, payload, lost in calls]
return {'id': name, 'results': results, 'reservations': len(ledger),
'reservedUnits': totals(ledger)}
def main:
base = {'target': 'prod-west', 'units': 500}
fixtures = [
fixture('acknowledged', [('reserve:prod-west:r70', base, False)]),
fixture('lost-ack-stable-key', [('reserve:prod-west:r70', base, True),
('reserve:prod-west:r70', base, False)]),
fixture('lost-ack-fresh-key', [('run-1', base, True), ('run-2', base, False)]),
fixture('changed-intent', [('op-1', base, True),
('op-1', {'target': 'prod-west', 'units': 700}, False)]),
fixture('wrong-target', [('op-1', base, True),
('op-1', {'target': 'prod-east', 'units': 500}, False)]),
fixture('two-authorized-operations', [('op-1', base, False),
('op-2', base, False)]),
]
by_id = {row['id']: row for row in fixtures}
assert by_id['lost-ack-stable-key']['reservedUnits'] == 500
assert by_id['lost-ack-stable-key']['results'][1]['status'] == 'replayed'
assert by_id['lost-ack-fresh-key']['reservedUnits'] == 1000
for name in ['changed-intent', 'wrong-target']:
assert by_id[name]['results'][1]['status'] == 'conflict'
assert by_id[name]['reservedUnits'] == 500
assert by_id['two-authorized-operations']['reservations'] == 2
replay_checks = 0
for units in range(1, 41):
for retries in range(1, 21):
ledger = {}
payload = {'target': 'prod-west', 'units': units}
reserve(ledger, 'stable', payload, lose_ack=True)
for _ in range(retries):
result = reserve(ledger, 'stable', payload)
assert result['status'] == 'replayed' and not result['created']
assert totals(ledger) == units and len(ledger) == 1
replay_checks += 1
conflict_checks = 0
for changed_units in range(1, 41):
ledger = {}
reserve(ledger, 'stable', base)
before = deepcopy(ledger)
assert reserve(ledger, 'stable', {'target': 'prod-west',
'units': changed_units})['status'] == 'conflict'
assert ledger == before
conflict_checks += 1
invalid = [('', base), ('x', None), ('x', {}),
('x', {'target': '', 'units': 5}),
*[('x', {'target': 'prod', 'units': n})
for n in [0, -1, True, 1.5, '5']]]
for key, payload in invalid:
ledger = {}
try:
reserve(ledger, key, payload)
except ValueError:
assert not ledger
else:
raise AssertionError('invalid input accepted')
payload = deepcopy(base); ledger = {}
reserve(ledger, 'stable', payload)
payload['units'] = 900
assert totals(ledger) == 500
print(json.dumps({'fixtures': fixtures, 'replayCombinations': replay_checks,
'conflictChecks': conflict_checks, 'invalidInputs': len(invalid),
'payloadCopyCheck': True, 'network': False, 'persistentWrites': False,
'vendorExecution': False, 'concurrencyProof': False,
'independentVerification': False,
'scriptSha256': sha256(Path(__file__).read_bytes).hexdigest}, indent=2))
if __name__ == '__main__':
main
r70 reserved 500 units before losing its ACK. Repeating the same key and payload returns the existing receipt; changing the ID creates another reservation. Moving to 700 under the old key requires resolving an intent conflict.
Common pitfalls
Confusing abandonment with cancellation; treating a lost ACK as no effect; using a fresh key on every retry; ignoring a freeze because the release is old; declaring recovery while convergence is still in progress.
Related topics: Pipeline verification and promotion · Incidents and RUN handover · Delivery identity and authorization
Decide from target state and confirmed effects. Preserve intent identity across retries and explicitly assign recovery ownership.
Reference: Roll back a target · Current linked guide; edition date unconfirmed (2026-09-30 inspection)