Define the service being handed to RUN
An APS team at a fictional bank rolls back a funds-service release during a maintenance window. The committee wants to know when operational ownership can return to RUN. This case does not describe internal BNP Paribas procedures. The decision requires more than an accepted command: identify the recovered revision, destination, functional tests and ability to detect the next incident. In Cloud Deploy, rollback creates a new rollout from an earlier release. The existence of that rollout identifies an execution to follow; an old rollout’s success does not establish the current execution’s result. Record both identifiers so a dashboard cannot present historical success as new evidence. Link each test to the revision and target actually used in the rehearsal. If the application depends on a data change incompatible with the earlier version, selecting that release does not resolve compatibility. This is an operational inference for the case rather than a guarantee supplied by the command. Give the compatibility check its own owner and record its result separately. For the guided exercise, write three acceptance conditions before rollback: the service processes a representative operation, applicable controls remain active and RUN receives the agreed failure signal. For each condition, decide who observes, which resource is observed and when evidence becomes too old. These deadlines belong to the exercise’s fictional contract. They are neither vendor defaults nor authorization to execute real changes.
Recover with the correct execution configuration
A release retains the pipeline and target instance that existed when it was created. The team later changes current target configuration and assumes rollback will automatically use that change. When a mismatch warning appears, confirming without reviewing the difference can execute a plan different from the expected one. Compare configuration associated with the release against current configuration and identify fields affecting destination, permissions and execution sequence. Another group suspended the pipeline to prevent additional changes during the incident. That measure also prevents rollback, redeployment and other documented operations. The recovery plan must include the suspension decision, responsible people and the point at which changes should be restricted again if needed. Do not assume that calling the situation an emergency creates an automatic exception. A control used to stop propagation of the problem can also block the chosen recovery mechanism. This interaction should be rehearsed before the incident window, with the relevant approval responsibilities recorded. In the guided case, the production owner presents a newly configured target, an older release and a suspended pipeline. Ask the learner to describe the analysis sequence: identify the intended release, review the mismatch, obtain the operational suspension decision and follow the new rollout. Avoid a universal command recipe for bypassing controls. The actual order depends on authorization and service design. The expected output is a reasoned decision, evidence of what will execute and a clear stopping point if the destination differs from the approved one.
Preserve artifacts required by rollback
A service can have a correct rollback procedure and still lose its required image through cleanup policy. The recovery inventory should connect each release to the artifact that must remain available. In Artifact Registry, a keep policy matching the same version as a delete policy preserves that version. Visual rule order must not be used to conclude the opposite. A cleanup rehearsal needs evidence of its own. Configuring dry run and observing no immediate events does not establish that the policy is harmless. Processing is periodic, and the service’s Data Write logs need configuration so results can be examined. Before enabling deletion, inspect predicted candidates, recovery artifacts and repository identity. Keep the decision to remove images separate from the promise to recover storage capacity by a particular deadline. Assign someone to check the completed dry-run evidence before making that promise. If the repository has immutable tags, tagged artifacts cannot be deleted under that configuration. A capacity estimate counting them as removable can therefore be wrong. For the exercise, present three versions: one matching only delete, another also matching keep and a tagged version subject to immutability. Ask for the expected outcome and evidence needed to confirm it. Then ask which versions support the rollback plan. Keeping an image does not prove it is approved, compatible or safe to return to production; it establishes only one recovery dependency.
Interpret maintenance by step and by VM
A patch job reaches the end of its window at two in the morning, but the team must plan follow-up for processes that can continue, including downloads and reboots. Do not present the window to the business as guaranteed instant interruption of all activity. Assign an owner to each VM still transitioning and explain the difference between starting no new work and completing work already underway. Pre-patch and post-patch scripts support application preparation and validation. When rebooting is necessary before patching begins, the pre-patch script runs before that reboot. A connection-draining step therefore needs review at that point in the lifecycle. For scripts stored as Cloud Storage objects, the reference includes the object generation, fixing the version used by the job. A filename alone does not establish that executed code matches reviewed code. Preserve this reference with the patch evidence so a later investigation can identify the actual script. Result interpretation also depends on the contract. ExecStepConfig permits allowedSuccessCodes; if the approved list includes zero and twelve, exit code twelve is not a failure merely because it is nonzero. However, accepting that code establishes neither whole-job success nor functional health. In the exercise, the script reports “draining already completed” through an accepted code. The learner must separate success of that step, patch installation, reboots and application testing. Expanding the list to hide a real error is not a correction of the operational problem.
Distinguish silence from restored observability
After rollback, the dashboard stops displaying alerts. Before interpreting that silence, establish whether current measurements exist and how the condition handles missing data. In a MetricThreshold condition with nonzero duration, EVALUATION_MISSING_DATA_INACTIVE can make the condition unmet when data stops. That is not a healthy measurement. The production owner should request both a functional observation and an observation of the telemetry path. A metric-absence policy does not automatically solve every startup with no data either. Its behavior requires a successful measurement before the monitored absence; a producer that has never emitted any point might not meet the expected condition. Include confirmation that the series exists and that labels match the recovered resource. Creating a metric definition or a dashboard name does not produce that evidence. Make the first successful observation an explicit rehearsal checkpoint, then record the subsequent loss of data separately. Time windows can also explain unexpected results. A threshold condition’s retest window resets when an aligned measurement no longer violates. Three violating minutes, one normal measurement and two more violating minutes do not constitute five continuous minutes. When a previously disabled policy is re-enabled, evaluation can use the recent window, including data from before re-enabling. Record alignment, duration and the enabling time when investigating an apparently immediate alert. The aim is to explain the actual signal, avoiding threshold changes made merely to turn the dashboard green.
Limit suppression and correlate the correct resource
During maintenance, a whole-policy snooze can close open alerts and prevent new alerts or notifications while active. Administrative closure does not repair the application. The plan should record suppression scope and expiry, alternative observation during the window and who confirms return to expected operation. Do not use closed-alert counts as resolved-incident counts without identifying why closure occurred. When the intention is to limit snoozing to the recovery instance, check the capability actually supported. In the API, a snooze with a filter applies to one policy and combines multiple labels using AND. A descriptive name does not restrict scope. The learner should compare authorized selection against the set selected by the filter before accepting the change. An overly broad rule can remove visibility from resources outside maintenance. Record that comparison with the change so the next shift understands which resources remain observable. Correlating conditions also requires care. AND can combine conditions met by different resources. If the requirement is to observe both conditions on the same VM, evaluate AND_WITH_MATCHING_RESOURCE and preserve necessary resource labels after aggregation. In the example, the reconciliation VM has processing delay and the distribution VM has authentication failures; combining those facts does not establish both failures on one VM. Draw a resource-by-condition matrix before selecting the combiner. At handover, also confirm receipt by the agreed recipient: a correctly evaluated condition does not, by itself, prove notification delivery.
Preserve logs with explicit outcome and coverage
The investigation needs to preserve a log window before source retention expires. Cloud Logging’s retroactive copy operation can route existing entries to Cloud Storage, but cannot recover entries whose retention has already expired. Identify source bucket, location, time filter, destination and why the selected window covers the incident. An empty destination’s existence does not establish evidence preservation. The request creates an operation whose outcome must be followed. A terminal state containing an error is not success. If writing fails, reconcile available destination content and gaps before declaring the archive complete. Copied count, filter scope and reading resulting objects answer different questions. A successful operation with the wrong filter still fails the investigation’s scope. Record the accepted scope alongside the operation identifier so another analyst can compare intent with results instead of relying on a summary status alone. Cancellation does not undo earlier copying either. Data already copied remains, and ongoing processes can finish before cancellation completes. In the exercise, a team cancels after choosing the wrong interval. Ask the learner to propose reconciliation: identify produced objects, separate correct from incorrect scope and define the next authorized operation. Avoid automatically deleting the destination to make it appear empty. Preservation, access, retention and eventual removal require separate decisions. The report should let another analyst reconstruct what was requested, what ended and what remains unproven.
Fresh-evidence exercise for handover
Run python3 run.py without credentials. The program uses fictional integer seconds, declared resources and three example checks: health, notification and rollback. It does not simulate Cloud Monitoring. For each resource and check, it selects the newest observation, requires the expected revision and checks that observation occurred at or after recovery and within the agreed maximum age. A recent failure supersedes an old success; the program does not select the last favorable result. If equally recent observations have the same timestamp but contradictory outcomes or revisions, the exercise requires reconciliation. An observation from another resource cannot satisfy a missing check. Suppression uses an interval closed at its start and open at its end, defined only for this model: at end it is no longer active. The code separates satisfied checks from complete inventory and keeps authenticity, actual notification and handover authorization claims false. Before running, predict outcomes for a wrong revision, old evidence, a missing test, an unknown result and active suppression. Then compare reasons and evidenceIds against your prediction. Tests check time boundaries, result combinations, input order and rejection of invalid fields. Supplied-data consistency does not establish that actual tests occurred. Finish by writing a RUN note with revision, scope, observation times, owners and gaps. Solve the final case by explaining which evidence would permit acceptance and which condition still prevents that decision.
"""Original offline handover worksheet, not a Cloud Monitoring emulator.
All timestamps are synthetic integer seconds. Freshness and required checks are
fictional acceptance rules, not vendor defaults. Evidence is not authenticated.
"""
from copy import deepcopy
from itertools import permutations, product
from hashlib import sha256
from pathlib import Path
import json
def label(x):
return isinstance(x,str) and bool(x.strip) and x==x.strip
def integer(x):
return type(x)is int and x>=0
def unique(xs):
return isinstance(xs,list) and bool(xs) and all(label(x)for x in xs) and len(xs)==len(set(xs))
def validate(m):
if not isinstance(m,dict) or set(m)!={'now','maxAge','inventoryComplete','resources','evidence','suppressions'}:
raise ValueError('Expected exact worksheet fields')
if not integer(m['now']) or not integer(m['maxAge']) or type(m['inventoryComplete'])is not bool:
raise ValueError('Integer clock, age and boolean inventory flag required')
if not isinstance(m['resources'],list) or not m['resources'] or not isinstance(m['evidence'],list) or not isinstance(m['suppressions'],list):
raise ValueError('Lists and nonempty resource scope required')
resources={}
for r in m['resources']:
if not isinstance(r,dict)or set(r)!={'id','revision','recoveryAt','checks'}:
raise ValueError('Expected exact resource fields')
if not label(r['id'])or r['id']in resources or not label(r['revision'])or not unique(r['checks'])or not integer(r['recoveryAt'])or r['recoveryAt']>m['now']:
raise ValueError('Invalid resource scope')
resources[r['id']]=r
ids=set
for e in m['evidence']:
if not isinstance(e,dict)or set(e)!={'id','resource','revision','check','observedAt','result'}:
raise ValueError('Expected exact evidence fields')
if not label(e['id'])or e['id']in ids or not label(e['revision'])or not label(e['resource'])or not label(e['check'])or e['resource']not in resources or e['check']not in resources[e['resource']]['checks']:
raise ValueError('Invalid or duplicate evidence identity')
if not integer(e['observedAt'])or e['observedAt']>m['now']or e['result']not in ['pass','fail','unknown']:
raise ValueError('Invalid evidence observation')
ids.add(e['id'])
ids=set
for s in m['suppressions']:
if not isinstance(s,dict)or set(s)!={'id','resource','check','start','end'}:
raise ValueError('Expected exact suppression fields')
if not label(s['id'])or s['id']in ids or not label(s['resource'])or not label(s['check'])or s['resource']not in resources or s['check']not in resources[s['resource']]['checks']:
raise ValueError('Invalid suppression scope')
if not integer(s['start'])or not integer(s['end'])or s['start']>=s['end']:
raise ValueError('Invalid half-open suppression interval')
ids.add(s['id'])
return resources
def analyze(m):
resources=validate(m);out=[]
for id,r in sorted(resources.items):
checks=[]
for check in sorted(r['checks']):
events=[e for e in m['evidence']if e['resource']==id and e['check']==check]
latest=max((e['observedAt']for e in events),default=None)
selected=[e for e in events if e['observedAt']==latest]
reasons=[]
if not selected:reasons.append('missing')
else:
# No arbitrary tie-break: conflicting newest observations need reconciliation.
if len({(e['revision'],e['result'])for e in selected})>1:reasons.append('conflicting-latest')
if any(e['revision']!=r['revision']for e in selected):reasons.append('revision-mismatch')
if latest<r['recoveryAt']:reasons.append('before-recovery')
if m['now']-latest>m['maxAge']:reasons.append('stale')
if any(e['result']=='fail'for e in selected):reasons.append('failed')
if any(e['result']=='unknown'for e in selected):reasons.append('unknown')
suppressed=sorted(s['id']for s in m['suppressions']if s['resource']==id and s['check']==check and s['start']<=m['now']<s['end'])
if suppressed:reasons.append('suppressed')
checks.append({'check':check,'latestObservedAt':latest,'evidenceIds':sorted(e['id']for e in selected),
'suppressionIds':suppressed,'reasons':sorted(reasons),'meetsDeclaredCheck':not reasons})
consistent=all(c['meetsDeclaredCheck']for c in checks)
out.append({'resource':id,'revision':r['revision'],'checks':checks,'allDeclaredChecksMet':consistent,
'inventoryCoverageUnproven':not m['inventoryComplete'],
'supportedUnderDeclaredEvidence':consistent and m['inventoryComplete'],
'evidenceAuthenticityVerified':False,'cloudPolicyBehaviorSimulated':False,
'realNotificationDelivered':False,'productionHandoverAuthorized':False})
return out
def model:
return {'now':1000,'maxAge':100,'inventoryComplete':True,
'resources':[{'id':'funds','revision':'r8','recoveryAt':850,'checks':['health','notification','rollback']}],
'evidence':[{'id':'e'+str(i),'resource':'funds','revision':'r8','check':c,'observedAt':950,'result':'pass'}for i,c in enumerate(['health','notification','rollback'])],
'suppressions':[]}
def evidence:
fixtures=[]
def record(id,m):
old=deepcopy(m);r=analyze(m);assert old==m
fixtures.append({'id':id,'results':r});return r
def reasons(r,check='health',resource=0):return next(c['reasons']for c in r[resource]['checks']if c['check']==check)
assert record('current-complete-evidence',model)[0]['supportedUnderDeclaredEvidence']
m=model;m['evidence'].pop(0);assert reasons(record('missing-check',m))==['missing']
m=model;m['evidence'][0]['observedAt']=899;assert reasons(record('stale-observation',m))==['stale']
m['evidence'][0]['observedAt']=900;assert record('exact-freshness-boundary',m)[0]['supportedUnderDeclaredEvidence']
m=model;m['maxAge']=200;m['evidence'][0]['observedAt']=849;assert reasons(record('before-recovery',m))==['before-recovery']
m['evidence'][0]['observedAt']=850;assert record('at-recovery-boundary',m)[0]['supportedUnderDeclaredEvidence']
m=model;e=deepcopy(m['evidence'][0]);e.update(id='new',observedAt=960,result='fail');m['evidence'].append(e);assert reasons(record('new-failure-supersedes-pass',m))==['failed']
m=model;m['evidence'][0]['result']='unknown'assert reasons(record('unknown-result',m))==['unknown']
m=model;m['evidence'][0]['revision']='r7'assert reasons(record('wrong-revision',m))==['revision-mismatch']
m=model;e=deepcopy(m['evidence'][0]);e.update(id='conflict',result='fail');m['evidence'].append(e);assert reasons(record('conflicting-latest',m))==['conflicting-latest','failed']
m=model;e=deepcopy(m['evidence'][0]);e['id']='corroborating'm['evidence'].append(e);assert record('consistent-latest-observations',m)[0]['supportedUnderDeclaredEvidence']
m=model;m['suppressions']=[dict(id='maintenance',resource='funds',check='notification',start=990,end=1010)];assert reasons(record('active-suppression',m),'notification')==['suppressed']
m['suppressions'][0]['end']=1000;assert record('at-suppression-end',m)[0]['supportedUnderDeclaredEvidence']
m=model;m['resources'].append(dict(id='payments',revision='r8',recoveryAt=850,checks=['health']));r=record('cannot-borrow-another-resource',m);assert r[0]['supportedUnderDeclaredEvidence']and not r[1]['supportedUnderDeclaredEvidence']
m=model;m['inventoryComplete']=False;r=record('incomplete-inventory',m)[0];assert r['allDeclaredChecksMet']and not r['supportedUnderDeclaredEvidence']
m=model;e=deepcopy(m['evidence'][0]);e.update(id='late-old-revision',revision='r7',observedAt=960);m['evidence'].append(e);assert reasons(record('newer-wrong-revision',m))==['revision-mismatch']
combinations=0
for results in product(['pass','fail','unknown'],repeat=3):
m=model
for e,r in zip(m['evidence'],results):e['result']=r
assert analyze(m)[0]['supportedUnderDeclaredEvidence']==all(r=='pass'for r in results);combinations+=1
timing=0
for at,age,suppressed in product([899,900,1000],[0,100],[False,True]):
m=model;m['maxAge']=age
for e in m['evidence']:e['observedAt']=at
if suppressed:m['suppressions']=[dict(id='maint',resource='funds',check='notification',start=999,end=1001)]
assert analyze(m)[0]['supportedUnderDeclaredEvidence']==(1000-at<=age and not suppressed);timing+=1
m=model;expected=analyze(m);orders=0
for order in permutations(m['evidence']):
v=deepcopy(m);v['evidence']=list(order);assert analyze(v)==expected;orders+=1
bad=[]
def altered(fn):
m=model;fn(m);bad.append(m)
altered(lambda m:m.update(now=True))
altered(lambda m:m.update(maxAge=-1))
altered(lambda m:m.update(inventoryComplete='yes'))
altered(lambda m:m.update(resources=[]))
altered(lambda m:m['resources'].append(deepcopy(m['resources'][0])))
altered(lambda m:m['resources'][0].update(checks=[]))
altered(lambda m:m['resources'][0].update(checks=['health','health']))
altered(lambda m:m['resources'][0].update(recoveryAt=1001))
altered(lambda m:m['resources'][0].update(revision=''))
altered(lambda m:m['evidence'].append(deepcopy(m['evidence'][0])))
altered(lambda m:m['evidence'][0].update(observedAt=1001))
altered(lambda m:m['evidence'][0].update(observedAt=-1))
altered(lambda m:m['evidence'][0].update(result='green'))
altered(lambda m:m['evidence'][0].update(resource='unknown'))
altered(lambda m:m['evidence'][0].update(resource=[]))
altered(lambda m:m['evidence'][0].update(check={}))
altered(lambda m:m['evidence'][0].update(check='unknown'))
altered(lambda m:m['evidence'][0].update(extra=True))
altered(lambda m:m.update(extra=True))
for start,end,resource,check in [(1000,1000,'funds','health'),(1001,1000,'funds','health'),(0,2,'unknown','health'),(0,2,'funds','unknown')]:
m=model;m['suppressions']=[dict(id='s',resource=resource,check=check,start=start,end=end)];bad.append(m)
bad.extend([None,[]])
for m in bad:
try:analyze(m)
except ValueError:pass
else:raise AssertionError('Invalid worksheet accepted')
return dict(scriptSha256=sha256(Path(__file__).read_bytes).hexdigest,fixtures=fixtures,
resultCombinations=combinations,timingCombinations=timing,inputPermutations=orders,invalidInputs=len(bad),
inputPreserved=True,orderIndependent=True,network=False,cloudExecuted=False,persistentWrites=False)
if __name__=='__main__':print(json.dumps(evidence,ensure_ascii=False,indent=2))
The current rollout exists, but the test belongs to the previous revision and snoozing closed the alerts. RUN handover still lacks sufficient evidence.
Common pitfalls
Counting closed alerts as resolved incidents; selecting the last favorable test; assuming suspension permits rollback; confusing a finished operation with success.
Related topics: Recover data, keys and AI controls · Failover with preserved capacity and trust · Secure automation and operational detection
RUN handover needs current observations of the correct revision and resource, including the ability to detect and communicate the next failure.
Reference: Roll back a target · Current linked guide; edition date unconfirmed (2026-09-30 inspection)