Define what still needs demonstration
An operational readiness review connects entry-to-service conditions to concrete observations. For a fictional funds-processing service, the list includes APS access, routing an alert, and executing a recovery task. Two positive results and one unexecuted task mean incomplete evidence. They establish neither failed recovery nor readiness by majority. Record the criterion, executor, version and configuration identity, environment, time, observed outcome, and limitation. The decision owner needs to understand the exact gap, its consequence, and who can resolve it. AWS readiness guidance describes a practice to adapt to context, rather than a mandatory bank form. This exercise also does not let an accepted exception become evidence that a task passed. If an explicit decision permits proceeding with an unmet condition, retain that condition separately from the decision and its boundaries. A list of green document links is not equivalent to these observations.
Rehearse with the future operator
Ask the APS operator to perform the task using their own account, the applicable procedure, and contacts available during the window. Watching the author use higher privileges does not establish operator autonomy. A runbook should identify prerequisites, steps, expected results, error handling, and escalation. If step four differs between the document and service state, pause that procedural handover and resolve the correct combination before treating it as usable. An R8/C3/E2 rehearsal does not automatically cover R8/C4/E2: here C4 changes the alert destination, so routing and receipt need checking in the target context. Earlier evidence remains useful for what it actually exercised. It need not be deleted, but cannot be extended to unobserved properties. Also establish that people required for recovery remain available after installation; attendance only at the start does not cover a late recovery decision. Record who can act if the first contact becomes unavailable.
Calculate the window with actual constraints
The teaching plan has an eight-minute drain, twelve-minute application work and eighteen-minute middleware work in parallel, a ten-minute smoke test, and fifteen-minute validation. Because the join waits for both branches, minimum duration is 8 + max(12,18) + 10 + 15 = 51 minutes. If both branches require the same executor and that person can perform only one task at a time, they become sequential: 63 minutes. For this exercise, the agreement also requires retaining 25 recovery minutes after normal execution. The requirement is therefore 88 minutes, 33 beyond the 55-minute window. Obtaining a second qualified executor reduces the requirement to 76 minutes, still leaving a 21-minute gap. This extra reserve is an explicit case constraint, not a universal rule or model of every recovery path. Before confirmation, revisit scope, sequence, resources, or the window with the responsible people. Opening two sessions does not remove the stated human constraint.
Distinguish observation, exception, and decision
The dashboard shows zero errors, but its last sample is six minutes old and the fictional criterion allows at most two minutes. The signal is stale: it establishes neither current health nor, alone, an outage. Likewise, urgent authority for incident X until 18:00 does not cover execution at 19:00 with feature Y added. Expose time and scope gaps to the defined authority. If an incident occurs, align recovery with incident coordination; two competing plans may change each other's state. Save the complete code below as release-readiness.py and run python3 release-readiness.py. Its fourteen checks exercise only arithmetic, declared dependencies, and gaps in supplied observations. Change a duration or observation and predict which assertion will fail before running it. The program does not measure a service, verify actual permissions, or authorize production. Submit a note identifying the gap, required evidence, owner, and decision time. Use that note to explain the decision with the specialists who own the affected tasks.
"""Original DR teaching fixtures. No network, deployment or production approval."""
import hashlib
import json
import platform
from pathlib import Path
def finish_times(tasks):
"""tasks: name -> (nonnegative minutes, already-declared predecessors)."""
finished = {}
for name, (minutes, predecessors) in tasks.items:
if minutes < 0 or any(p not in finished for p in predecessors):
raise ValueError("Durations must be nonnegative and predecessors declared first")
finished[name] = max((finished[p] for p in predecessors), default=0) + minutes
return finished
def requirement_gaps(required, observations):
# A missing result is different from an observed failure.
return {name: observations.get(name, "not-run") for name in required
if observations.get(name)!= "pass"}
def changed_context(observed, target):
return sorted(k for k in observed.keys | target.keys
if k not in observed or k not in target or observed[k]!= target[k])
def sample_state(now_minute, sampled_minute, maximum_age):
age = now_minute - sampled_minute
if age < 0:
return "clock-inconsistent"
return "current" if age <= maximum_age else "stale"
def exception_gaps(now_minute, expires_minute, requested, covered):
# This exercise makes expiry exclusive. Real policy must define its own boundary.
return {"expired": now_minute >= expires_minute,
"uncovered": sorted(set(requested) - set(covered))}
def run:
checks = []
def check(name, actual, expected):
if actual!= expected:
raise AssertionError((name, actual, expected))
checks.append({"name": name, "actual": actual, "passed": True})
parallel = {
"drain": (8, []), "app": (12, ["drain"]),
"middleware": (18, ["drain"]),
"smoke": (10, ["app", "middleware"]), "validate": (15, ["smoke"]),
}
serial = {**parallel, "middleware": (18, ["app"])}
p, s = finish_times(parallel), finish_times(serial)
check("parallel_dependency_timing", p,
{"drain": 8, "app": 20, "middleware": 26, "smoke": 36, "validate": 51})
check("exclusive_executor_timing", s,
{"drain": 8, "app": 20, "middleware": 38, "smoke": 48, "validate": 63})
window, reserve = 55, 25
check("window_plus_explicit_recovery_reserve",
{"parallelRequired": p["validate"] + reserve,
"serialRequired": s["validate"] + reserve,
"parallelGap": p["validate"] + reserve - window,
"serialGap": s["validate"] + reserve - window},
{"parallelRequired": 76, "serialRequired": 88, "parallelGap": 21, "serialGap": 33})
required = ["aps-access", "alert-route", "recovery-task"]
observations = {"aps-access": "pass", "alert-route": "pass"}
check("missing_is_not_observed_failure", requirement_gaps(required, observations),
{"recovery-task": "not-run"})
check("failed_is_retained", requirement_gaps(required, {**observations, "recovery-task": "fail"}),
{"recovery-task": "fail"})
check("all_observed_is_only_model_coverage", requirement_gaps(required, {**observations, "recovery-task": "pass"}), {})
observed = {"build": "R8", "config": "C3", "environment": "E2"}
check("same_build_different_config", changed_context(observed, {**observed, "config": "C4"}), ["config"])
check("absent_target_context_is_not_equivalence", changed_context(observed, {"build": "R8", "config": "C3"}), ["environment"])
check("stale_zero_error_sample", sample_state(19 * 60, 19 * 60 - 6, 2), "stale")
check("freshness_boundaries", [sample_state(100, 98, 2), sample_state(100, 101, 2)], ["current", "clock-inconsistent"])
check("exception_time_and_scope_are_separate", exception_gaps(19 * 60, 18 * 60, ["incident-X", "feature-Y"], ["incident-X"]),
{"expired": True, "uncovered": ["feature-Y"]})
check("exception_expiry_boundary", [exception_gaps(t, 1080, ["incident-X"], ["incident-X"])["expired"] for t in [1079, 1080]], [False, True])
milestones = {"apiAvailable": 20 * 60 + 10, "backlogComplete": 20 * 60 + 45, "reconciled": 21 * 60}
check("distinct_recovery_milestones",
[milestones["backlogComplete"] - milestones["apiAvailable"], milestones["reconciled"] - milestones["backlogComplete"], milestones["reconciled"] - milestones["apiAvailable"]], [35, 15, 50])
residual = {"status": "accepted-open", "owner": "service-owner", "reviewDay": 10}
check("accepted_is_not_closed_and_review_can_be_overdue", {"status": residual["status"], "daysOverdue": max(0, 12 - residual["reviewDay"])},
{"status": "accepted-open", "daysOverdue": 2})
return {"scope": "Synthetic arithmetic and evidence-gap fixtures only. No production gate, deployment, service health measurement or authorization decision. Durations and thresholds are fictional; the input observations are assumed, not collected.",
"python": platform.python_version, "scriptSha256": hashlib.sha256(Path(__file__).read_bytes).hexdigest,
"groups": len(checks), "checks": checks}
if __name__ == "__main__":
print(json.dumps(run, indent=2, ensure_ascii=False))A 55-minute window, one executor, and an extra 25-minute reserve: execution takes 63, total required is 88, and the gap is 33 minutes. A second person reduces the gap to 21 without removing the need for a decision.
Common pitfalls
Confusing unexecuted with failed; accepting zero errors from stale samples; treating branches as parallel despite an exclusive executor; extending an exception without a new decision.
Related topics: Capacity and dependencies · Operational readiness · Recovery and escalation
Establish evidence applicability and sequence feasibility before requesting the window decision.
Reference: OPS07-BP02 Ensure a consistent review of operational readiness · Google SRE release and canary guidance; GitHub immutable releases and GitLab release evidence and deployment safety; DORA five-metric model; inspected 2026-10-01