1. Define the recovery outcome
Consider a fictional service receiving instructions for a funds batch. Following an incident, the application responds again and its queue shrinks. The business owner still needs to know which instructions produced valid outcomes. Define three separate observations: accepted work, pending work and confirmed outcomes. A processing attempt can fail; a defect can acknowledge work before a correct write; quarantine forwarding can remove a message from the active queue. None of these events alone establishes that the batch completed. Before intervention, write the acceptance criterion. In this exercise, every expected identifier must have exactly one valid outcome or an exception explicitly accepted by its owner. This is a teaching condition for evaluating decisions, not a procedure attributed to a bank or a universal rule. Define who reconciles outcomes, who may accept exceptions and who owns work remaining after the window. Infrastructure confirms capacity and availability; the application explains produced effects; the batch owner decides acceptance using agreed evidence. A green dashboard state does not replace these responsibilities. Include the identifiers and evidence locations in the handover so that the next team can continue the investigation without rebuilding the history.
2. Calculate the headroom that reduces backlog
In the simplest model, every operation costs the same, failures do not occur and rates stay constant. If 80 operations arrive each second and 120 complete, the queue loses 40 each second. A backlog of 12,000 takes 300 seconds to clear. Dividing only by 120 would give 100 seconds and ignore work still arriving. With 90 arrivals and 90 completions per second, an existing backlog stays constant. Restoring normal capacity can stabilize service without recovering accumulated work. Calculate using effective completions. If workers support 200 per second but the database caps completions at 110, useful capacity in this model is 110. With arrivals of 80 and 9,000 pending, recovery needs five minutes. Doubling workers without changing the downstream cap does not alter the calculation. In practice, measure intervals and track changes in operation cost. A migration can introduce more expensive queries, and replay can have a different workload mix from current traffic. Present the forecast with its assumptions, observation time and next review. The formula supports reasoning about a controlled scenario; it does not turn nominal capacity into a deadline promise.
3. Prepare replay with explicit boundaries
To replay previously acknowledged Pub/Sub messages by timestamp, confirm applicable topic or subscription retention. Check that the required interval remains covered. A successful seek does not create an instantaneous boundary for every delivery; transition is eventually consistent. Evaluate recovered work against business criteria, rather than assuming queue position matches database state. Seek does not undo external writes. In the exercise, a faulty version acknowledged 100 instructions but confirmed only 70 outcomes. First preserve evidence and identify the 100 expected IDs. Then separate confirmed, missing and ambiguous outcomes. A write whose client timed out needs investigation: timeout does not prove that the transaction failed. Fix the consumer before exposing it to the batch again. Rehearsing a subset allows effects to be compared without immediately sending the entire volume to the dependency. Define stop conditions and who may resume. Record the corrected version, time scope and origin of the IDs. If recovery scope expands, make that change visible to the owner because it changes volume, risk and timing. Preserve the initial scope as a comparison point; silently replacing it makes later reconciliation harder.
4. Separate delivery and operation identity
Pub/Sub exactly-once is a pull-subscription feature, including StreamingPull, with a regional guarantee. It does not add this support to a push subscription. Even with the feature, two publishes can correspond to one business intent. A different message identifier does not establish that two operations were requested. The application needs to relate transport identity to the identity it uses to recognize business work. In a fictional example, instruction F-42 is published twice after an uncertain response. A consumer finds two message IDs and two records with the same amount. Investigate instruction origin, outcomes and the rule distinguishing a legitimate correction from a repeat. Do not delete a record merely because values match: distinct operations can also share an amount. Define stable identity for intent and treat a payload mismatch as an explicit condition. The earlier hook-recovery exercise illustrates an in-memory association; a real implementation needs persistence and concurrency handling. The objective here is to choose evidence that supports a decision, without assuming a transport setting solves every financial duplication. Keep ambiguous outcomes open for investigation instead of treating absent client confirmation as proof that no effect occurred.
5. Protect sequence and consumer capacity
Ordered callbacks per key do not automatically control asynchronous work that continues after a callback returns. If revisions 11 and 12 start parallel writes, 11 can finish later and overwrite newer state. The application must preserve sequence where effects occur. A global key can also concentrate independent work. If each portfolio requires internal order but different portfolios are independent, a stable key per portfolio is an option to assess. Transition needs to account for messages already in flight. Ask the business question first: which operations may genuinely change order? An apparent throughput increase can be invalid if it changes required sequencing. In rehearsal, compare each portfolio’s final outcome and include operations with different execution times. Random key distribution can hide the queue while producing incorrect results. The project team should allow time to confirm these invariants with the application; support should be able to locate the affected key or operation. When the problem occurs before publication, distinguish it from consumption: publisher flow control bounds outstanding messages and bytes in the client. That control does not establish downstream batch recovery. Record the boundary affected by each change so that separate teams do not apply unrelated settings to the same symptom.
6. Follow up quarantine and retries
Dead-letter forwarding is best-effort and the configured attempt count is approximate. Check configuration and service-identity permissions; an operator’s access to the topic does not replace those permissions. Forwarded messages need an analysis and recovery path. Assign ownership of quarantined work and criteria for correcting, replaying or accepting an exception. Avoid using the original queue becoming empty as the sole success criterion. In a closed batch of 1,000 operations, 980 completed and 20 quarantined still leave 20 outcomes to resolve. Do not shrink the denominator to present 100% without an explicit agreement changing acceptance. Retry has its own scope too: backoff applies per message and does not establish a global subscription pause. During dependency maintenance, the team needs an explicit plan for new work and work already delivered. Describe where each set remains preserved, how it will resume and who confirms recovery. A poorly defined pause can merely move accumulation into another process’s memory. Observe the boundary where work actually waits before changing limits. Include exception age and next action in the operational handover so that quarantine remains managed work instead of an invisible endpoint.
7. Rehearse decisions and communicate remaining work
Prepare three variants of one incident: stable arrivals, growing arrivals and reduced downstream capacity. Ask the learner for a forecast and the condition that invalidates it. With 9,000 pending, 80 arrivals/s and 110 completions/s, four minutes are insufficient. The decision may require controlled admission without losing instructions, effective added capacity or a renegotiated deadline. None should be presented as guaranteed without confirming dependencies and authorization from the service owner. When communicating, separate service availability, estimated backlog and batch acceptance. A useful update identifies measurement time, affected set, conditional forecast, exceptions and next update. At shift handover, provide IDs or searchable references, owners and closure criteria. A list of healthy servers does not explain that twelve operations still lack outcomes. Before closure, compare sets of identities. If A,B,C,D are expected but the journal contains A,A,B,C, four rows do not establish four distinct outcomes. A repeat needs investigation and D remains missing. Retaining that difference makes the next action concrete and prevents aggregate indicators from hiding incomplete work. Keep the case open until the agreed acceptance evidence exists or an authorized exception changes its disposition.
8. Local work-conservation laboratory
Run the Python code shown below. The model uses integer seconds: it first adds arrivals, then completes the minimum of available work, worker capacity and downstream capacity. Operations have equal cost. An explicit number can move from the initial set into quarantine, which never counts as completed. Each step checks accepted = completed + pending + quarantined. There are no cloud calls, persistence, retries, expiry or simulation of Pub/Sub guarantees. Compare net-capacity and balanced. Then compare downstream-cap, which keeps 1,800 pending after 240 seconds, with downstream-drained, which reaches zero at 300. In load-change, the queue moves from 1,000 to 600 and ends at 700. In quarantine, the queue reaches zero but only 980 of 1,000 operations completed. Explain empty-then-regrows too: firstEmptySecond records the first visit to zero, not a promise of permanent recovery. The model checks 1,600 constant-rate combinations against an independent formula and rejects negative inputs, booleans and incomplete structures. Change a phase, predict the result by hand and only then execute. Evidence validates this exercise’s arithmetic, rather than real performance, message ordering or compliance with an SLA. Use the differences to explain why operational acceptance requires evidence beyond a simulated queue length.
"""Original teaching model. No Pub/Sub client, real-time scheduler or cloud calls."""
import hashlib
import itertools
import json
from pathlib import Path
def integer(value, name):
if type(value) is not int or value < 0:
raise ValueError(name + ' must be a nonnegative integer')
return value
def model(initial, phases, quarantined=0):
"""Each second admits arrivals, then completes min(queue, worker, downstream).
Initial quarantined work leaves the active queue but is never completed.
All units are equally costly. No failures, retries, latency or expiry.
"""
integer(initial, 'initial')
integer(quarantined, 'quarantined')
if quarantined > initial:
raise ValueError('quarantine exceeds initial work')
if not isinstance(phases, list):
raise ValueError('phases must be a list')
checked = []
for phase in phases:
if not isinstance(phase, dict) or set(phase)!= {'seconds', 'arrivals', 'worker', 'downstream'}:
raise ValueError('phase fields must match the model')
checked.append({k: integer(v, k) for k, v in phase.items})
queue, accepted, completed, elapsed = initial - quarantined, initial, 0, 0
first_empty = 0 if queue == 0 else None
points = []
for phase in checked:
capacity = min(phase['worker'], phase['downstream'])
for _ in range(phase['seconds']):
queue += phase['arrivals']
accepted += phase['arrivals']
done = min(queue, capacity)
queue -= done
completed += done
elapsed += 1
if queue == 0 and first_empty is None:
first_empty = elapsed
assert accepted == completed + queue + quarantined
points.append({'second': elapsed, 'pending': queue, 'completed': completed})
return {'accepted': accepted, 'completed': completed, 'pending': queue,
'quarantined': quarantined, 'firstEmptySecond': first_empty,
'allAcceptedCompleted': queue == 0 and quarantined == 0,
'phaseEnds': points}
def phase(seconds, arrivals, worker, downstream):
return dict(seconds=seconds, arrivals=arrivals, worker=worker, downstream=downstream)
def main:
scenarios = [
('net-capacity', 12000, [phase(300, 80, 120, 120)], 0, 0, 300),
('balanced', 4000, [phase(300, 90, 90, 90)], 0, 4000, None),
('downstream-cap', 9000, [phase(240, 80, 200, 110)], 0, 1800, None),
('downstream-drained', 9000, [phase(300, 80, 200, 110)], 0, 0, 300),
('load-change', 1000, [phase(10, 20, 60, 60), phase(10, 70, 60, 60)], 0, 700, None),
('quarantine', 1000, [phase(10, 0, 100, 100)], 20, 0, 10),
('empty-then-regrows', 10, [phase(1, 0, 10, 10), phase(2, 20, 10, 10)], 0, 20, 1),
('no-capacity', 4, [phase(3, 2, 0, 10)], 0, 10, None),
]
fixtures = []
for name, initial, phases, quarantine, pending, first_empty in scenarios:
before = json.dumps(phases, sort_keys=True)
result = model(initial, phases, quarantine)
assert result['pending'] == pending and result['firstEmptySecond'] == first_empty
assert json.dumps(phases, sort_keys=True) == before
fixtures.append({'id': name, **result})
assert fixtures[5]['completed'] == 980 and not fixtures[5]['allAcceptedCompleted']
assert fixtures[6]['firstEmptySecond'] == 1 and fixtures[6]['pending'] == 20
combinations = 0
for initial, arrivals, worker, downstream, seconds in itertools.product(range(5), range(4), range(4), range(4), range(5)):
result = model(initial, [phase(seconds, arrivals, worker, downstream)])
expected = max(0, initial + seconds * (arrivals - min(worker, downstream)))
assert result['pending'] == expected
assert result['accepted'] == initial + seconds * arrivals
assert result['completed'] == result['accepted'] - expected
combinations += 1
invalid = [(-1, [], 0), (True, [], 0), (1.5, [], 0), (1, [], -1), (1, [], 2),
(1, {}, 0), (1, [{}], 0), (1, [phase(1, -1, 1, 1)], 0),
(1, [phase(True, 1, 1, 1)], 0), (1, [phase(1, 1, '2', 1)], 0)]
for args in invalid:
try:
model(*args)
except ValueError:
pass
else:
raise AssertionError('invalid input accepted')
print(json.dumps({'scriptSha256': hashlib.sha256(Path(__file__).read_bytes).hexdigest,
'fixtures': fixtures, 'constantRateCombinations': combinations, 'invalidInputs': len(invalid),
'inputPreserved': True, 'cloudExecuted': False, 'network': False, 'persistentWrites': False,
'limitations': 'Integer equal-cost work model; no vendor execution, deadlines, retries, expiry, per-message order or delivery guarantees.'}, indent=2))
if __name__ == '__main__':
main
9,000 pending, 80 arrivals/s and an effective cap of 110 completions/s require five minutes in the constant-rate model.
Common pitfalls
Ignoring arrivals, counting attempts as completions, confusing seek with data rollback and hiding quarantine behind an empty queue.
Related topics: SLOs and service capacity · Hook recovery and intent identity · Telemetry during rollout
Recovery needs net capacity and evidence of expected outcomes, with exceptions tracked through acceptance.
Reference: Professional Cloud DevOps Engineer exam guide · Current linked guide; edition date unconfirmed (2026-09-30 inspection)