Publish and observe a second identity
On the actual NFS mount, the synthetic producer writes delivery.tmp, applies the intended mode and renames it to delivery.ready. The original name disappears and the final name becomes available. A second identity reads 23 bytes and confirms the expected hash. The script adds retry.ready containing the same instruction and counts two files for one identifier. No business consumer is implemented: the result does not say how many financial operations would occur. The workshop asks the learner to draw the boundary between publication, reading and processing. Also define how the consumer recognizes repeated intent and what happens when the same identifier returns with different content. A new filename alone does not resolve that contract.
Reason about a lost acknowledgement
The missing-acknowledgement case is a guided discussion, not an injected lab fault. After an NFS error, do not assume rename had no effect: compare names and content and query consumer state before retrying. Preserve intent identity to relate attempts. If the query remains pending, communicate uncertainty with an owner and review deadline instead of inventing success or failure. Another situation changes the publication mechanism: temporary files on another filesystem can make rename return EXDEV. Replacing the operation with direct copying to the final name requires reviewing when the consumer may begin reading. These decisions belong to the application contract and should be exercised with that application.
Measure recovery through data and useful service
In a fictional example, the incident occurs at 12:00 and the last confirmed recoverable point is 11:40. A ten-minute RPO is not demonstrated by that twenty-minute-old point even if restoration is fast. RTO evaluates another commitment: time until useful service under agreed criteria. The workshop uses a timeline containing incident, file restoration, reconciliation and consumer acceptance. If application acknowledgements are ahead of restored files, compare identifiers before replay. The RAM-backed lab does not test persistence after crashes or measure these objectives. Scenario times are fictional inputs for decision practice and require their own evidence in a representative recovery exercise.
Decide cutover and rollback with a visible delta
The final scenario has 18 deliveries created on the new target after cutover, seven of which have acknowledgements. Before returning to the source, contain writers and consumers under the approved plan, preserve inventory and classify each identity as confirmed, unprocessed or uncertain according to evidence. Restoring DNS neither removes business effects nor adds new data to the source. The team needs to agree the restart point and reconciliation owners. In the English RUN handover, include observed results, what remains unproven, rollback criteria and who authorizes reopening. An active mount and matching hashes are useful parts of the evidence but do not close outstanding functional or recovery criteria.
# Read recorded evidence, without starting services or changing mounts.
python3 - <<'PY'
import json
from pathlib import Path
r=json.loads(Path("content/labs/nas-nfs/evidence.json").read_text)
for key in ["publishByRename", "consumerReadsPublishedBytes",
"twoNamesOneInstruction", "unmountRevealsLocalDirectory"]:
print(key, r["checks"][key])
PYTwo deliveries have different filenames but the same instruction. In a fictional cutover, 18 new deliveries and seven acknowledgements require reconciliation before rollback.
Common pitfalls
Using hashes as processing proof, retrying with new identity, confusing RPO with RTO or reverting an endpoint while ignoring new data.
Related topics: Idempotency and uncertain outcomes · Disaster recovery and RUN handover
A delivery is accepted only within demonstrated scope. Recovery and rollback need reconciliation of data and consumer effects.
Reference: Disaster Recovery of Workloads on AWS: introduction · BigSavant NAS 2026-09; selected Linux NFS, Samba, Windows SMB and ONTAP behavior