Readiness does not cover every population
The faulty candidate parses configuration, announces ready and returns 22 for ordinary. The same calculation should return 22 for critical, but a deliberate condition adds one unit and produces 23. There is no parsing error or stopped process: a functional outcome is violated. Preflight and the ordinary sample do not detect that condition. Before expanding a fictional banking-processing change, define relevant populations and consumer criteria. Low volume does not waive a mandatory route when its consequence matters. The lab measures no real failure rate and establishes no canary sample size; it shows why the observed request selection limits the conclusion that can be communicated.
Recover processes and future outcomes
Returning current to v1 does not change configuration in the still-running v2-bug instance, which continues returning 23. Only after starting a v1 instance does the script observe 12 again. This sequence should not be reduced to a checked rollback box. The plan must select a recoverable version, define activation, manage instances and observe the path that will receive requests. The lab sends directly through pipes and exercises no load balancer, connection draining or service manager. At a real target, those mechanisms require their own rehearsal. Retain the recovery dependency until a valid retirement decision exists; reclaiming space by deleting the old version can remove an option the plan still requires.
Prior effects retain their own history
Before recovery, the faulty worker writes a synthetic receipt with result 23 to a separate file. Pointer replacement and v1 startup do not alter that record. The script reads it again and confirms it remains unchanged. The receipt represents an effect boundary; it is neither a payment nor a sent message. In a real application, identify affected requests, consumers, confirmed state and authorized reconciliation or compensation actions. Do not delete evidence to make history resemble current state. Technical closure can coexist with pending reconciliation only when criteria and owners explicitly permit it. Communication must distinguish recovered service, assessed effects and consequences still requiring treatment.
Operational acceptance with practicable decisions
The proposed workshop brings together APS, development, supplier and service owner. Participants should explain each instance’s state, the critical failure and persistent receipt in English, proposing criteria to expand, recover or pause. The RUN plan should include commands, applicability conditions, expected observations, stopping and escalation. File access does not establish autonomy; a written guide does not establish an executed session. Thirteen local groups passed twice, but the human workshop and independent review remain outstanding. In subsequent review, separate prevention, detection and recovery: a delivered alert implements neither preflight nor critical-route coverage. Each action needs an owner and a checkable outcome linked to the gap it is intended to resolve.
python3 content/labs/change-runtime/run.py --output /tmp/dr-change-recovery.json
# Compare readinessAndOneRouteMissDefect and rollbackDoesNotErasePriorEffect.
# Proposed workshop: explain technical recovery versus reconciliation in English.The candidate passes ordinary with 22, but critical returns 23. After recovery, new requests work and the incorrect receipt remains recorded.
Common pitfalls
Confusing readiness with correct outcomes, expanding from the aggregate or declaring every effect resolved after returning to earlier code.
Related topics: Canary and acceptance criteria · Reconciliation and learning
Technical recovery and reconciliation have distinct criteria; closure must show the state of both and ownership of outstanding work.
Reference: Canarying Releases · BigSavant Change Management 2026.1; independent technical curriculum