Distinguish revision from business time
A revision identifies the logical order of changes in etcd; it is neither a business timestamp nor a count of seconds. In the lab, snapshot revision is below the revision observed after additional writes. Ordinary restore returns to the captured point. A read explicitly requesting the higher source revision receives a future-revision rejection at the destination. Record all three values: snapshot revision, later observation, and restored revision. Their difference explains request behavior but does not calculate RPO in minutes. That calculation requires linking recoverable data to service times and operations.
Increase revision without inventing data
The second restore uses the same artifact, bump-revision=1000, and mark-compacted. Revision rises above the highest observation in the small test fixture, while values remain at the snapshot point. The subsequently created key stays absent. The increment was chosen for this bounded exercise and should not be copied as a production constant. In a real plan, estimate possible advancement since capture, consider snapshot age, and justify margin. If captured revision is 400 and the highest observation is 950, an increment of 551 exceeds that value by one; this is only the arithmetic boundary, without operational margin.
Interpret compaction and rebuild consumers
On the adjusted destination, the lab requests an old revision through both a historical read and a watch. Both report compaction. That signal demonstrates that requested history is unavailable; it does not demonstrate that an application rebuilt its cache. A consumer needs to handle the result according to its contract: obtain consistent state, reconcile additions and deletions, and resume watching at an appropriate position. The exercise runs no Kubernetes informers and does not validate that process for any framework. For application acceptance, compare rebuilt state and introduce a controlled subsequent change to observe continuity.
Observe continuation and preserve the artifact
After the bumped restore, the script writes another test key. It checks that revision advances and that all three members return that value. It also confirms that the original snapshot digest has not changed. These are different checks: the first observes progress, the second agreement on a bounded read, and the third artifact preservation. None proves recovery of missing later data. Watch guarantees also establish no universal upper bound on delivery latency. If a consumer has a lag objective, define measurement and exercise representative conditions, retaining the distinction between an API guarantee and a service objective.
Measure RPO and RTO against agreed function
Consider a fictional incident at 09:00, restore-command completion at 09:07, and function acceptance at 09:24. If agreed RTO is twenty minutes and function becomes available only at acceptance, recovery takes twenty-four minutes. Record the gap and intermediate timings to locate improvements. For RPO, use the last recoverable business point, including dependencies and additional sources actually available. Increasing revision does not make that point more recent. Do not change the objective after seeing the result; an approved exception may allow temporary operation, but it should preserve the original measurement and accepted risks.
Decision and English communication workshop
Allow forty-five minutes: ten to predict results, fifteen to run and compare evidence, ten to decide resumption, and ten to present status in English. Assign APS, application-owner, and sponsor roles. Use a case where the cluster recovered, a cache still contains an old rule, and the scheduler points to previous endpoints. Deliver a decision containing passed evidence, missing conditions, an owner, and a reassessment deadline. In the briefing, distinguish local results, operational hypotheses, and approval. The lab uses up to three simultaneous processes on one Darwin ARM64 host, listed as tier 3 in consulted documentation; it does not qualify a production architecture.
# After running the self-contained lab from the preceding lesson:
python3 - <<'PYCODE'
import json
from pathlib import Path
result = json.loads(Path('evidence.json').read_text)
assert result['passed'] == 10 and result['failed'] == 0
for check in result['checks']:
print(check['name'], check['observations'])
# This reads local exercise evidence. It does not query a production service.
PYCODEThe recovered cluster shows a higher revision while the cache remains stale. The team holds the consumer, reconciles state, and checks a subsequent change before accepting resumption.
Common pitfalls
Treating a bump as data recovery; repeating compacted cursors without handling the result; measuring RTO only by the command; attributing behavior to an informer the lab did not execute.
Related topics: Objectives and dependencies · Restore: recovered point and integrity · Resumption: replay and operational acceptance
Revision, contents, and consumer state each need their own evidence. Recovery reaches the agreed function with a known data point and explicit decisions about resumption risks.
Reference: Disaster recovery · BigSavant recovery 2026-09; PostgreSQL 18, etcd 3.6 and selected AWS/Azure behavior