← LFCE: Linux engineering and recovery
01 / 8 · 45 MIN

Automate with evidence

Control change scope and interpret simulation limits.

Concept and mechanism

A repeatable change needs an origin, scope, and success criterion. Git relates configuration to a reviewed decision. On an already shared branch, revert creates a new commit undoing a change without deleting the published sequence. This does not guarantee the resulting state is operationally correct: review and validation remain necessary. For file transformations, separate candidate generation from original replacement. A validated temporary file reduces the risk of publishing a partial write. Confirm permissions, ownership, and replacement behavior; do not assume universal atomicity across filesystems or tools.

Guided application

In a fictional batch-host rollout, begin with a small batch and include functional checks before advancing. In Ansible, serial controls play batches; forks controls workers and free permits independent host progression. Check mode can skip tasks or fail to reproduce dependencies, so a clean result does not demonstrate every path. If the first host fails after a kernel update, preserve the others and investigate early boot. For slowness, relate symptoms to resources: I/O pressure indicates waiting rather than proving which disk failed. The change record should retain the hypothesis, observations, and stopping condition so another colleague can continue diagnosis.

IN PRACTICE

A one-host batch is useful only if its validation prevents progression on failure.

Common pitfalls

Simulation as complete proof; uniformity before health; direct writes without a validated candidate.

Related topics: Operate and recover systems · Identity, PAM, and SSH · Networking, tunnels, and policies

Take this idea with you

Automate observation and the decision to stop as well.

Create account

Reference: Ansible execution strategies · Historical LFCE V3.18 (2018-06-12); exam retired 2022-05-01; technical references inspected 2026-09-30