Recovery is not approval
During an overnight window, the new worker starts but produces a reconciliation result incompatible with acceptance criteria. The team can return to previous configuration. There are now two conclusions: recovery may have succeeded and the release remains rejected. If recovery ends in a successful rescue without another signal, a consumer looking only at exit code may interpret execution as an accepted change. Define the contract with the pipeline and incident manager in advance. In this lab, a fail task after restoration keeps the overall rejection result.
Read, change, fail and restore
recovery.yml reads old bytes with slurp, writes different synthetic configuration and injects an acceptance failure. rescue uses saved content to restore the file with mode 0600. The next task preserves the release error. always records that the recovery path was visited. The runner requires exit code 2, compares final bytes with initial bytes and looks for that record on both workers. These checks avoid calling the mere presence of a rollback message a success.
What the mechanism does not resolve
An unreachable machine is not equivalent to a task returning failed. A playbook definition error can also prevent the expected flow. Therefore, do not treat always as a promise of recording under every circumstance. Plan external execution evidence and escalation when a node stops responding. Content saved by slurp exists only in this run; recovery after controller loss would require suitable persistent artifacts and procedures. If restoration fails, the final message must retain that uncertainty rather than announce a healthy service.
Criteria for returning to operations
In the example, matching hashes prove restoration of synthetic files. They do not prove that a process reloaded configuration, queues were reconciled or an external transaction was reversed. In a real application, add specific probes and assign responsibility for reopening traffic. Record before and after versions, initial failure, executed actions and residual risks. A useful report distinguishes rejected change with recovered service, rejected change with partial recovery and unknown state. These distinctions help APS and project management communicate the same outcome.
- name: Restore a local configuration but reject the release
hosts: workers
gather_facts: false
vars:
node_dir: "{{ lab_root }}/{{ inventory_hostname }}"
tasks:
- name: Read known previous bytes
ansible.builtin.slurp:
src: "{{ node_dir }}/worker.conf"
register: previous
- name: Exercise an unsuccessful rollout
block:
- name: Simulate accepted syntax but failed business behavior
ansible.builtin.copy:
content: "worker={{ inventory_hostname }}\nport=9443\n"
dest: "{{ node_dir }}/worker.conf"
mode: '0600'
- name: Inject a synthetic acceptance failure
ansible.builtin.fail:
msg: Synthetic batch acceptance failed
rescue:
- name: Restore the previous configuration bytes
ansible.builtin.copy:
content: "{{ previous.content | b64decode }}"
dest: "{{ node_dir }}/worker.conf"
mode: '0600'
- name: Keep the release gate failed after recovery
ansible.builtin.fail:
msg: Configuration restored; release still rejected
always:
- name: Record that recovery was attempted
ansible.builtin.copy:
content: "recovery-path-visited\n"
dest: "{{ node_dir }}/recovery.log"
mode: '0600'The file returns to 8443 and execution ends with failure: file recovery and release rejection are compatible outcomes.
Common pitfalls
Rescue as acceptance; always as universal guarantee; configuration hash as business probe.
Related topics: Executable rehearsal: scope, state and limits · Validated configuration and observable activation · Archives, residual files and markers
Measure recovery and preserve the change outcome in the execution contract.
Reference: Block error handling · EX294 current objectives inspected 2026-09-30; RHCE in Ansible framework effective 2026-05-11; booking product version not publicly pinned