← Disaster Recovery: prepare, recover, and validate
09 / 12 · 60 MIN

Snapshot, identity, and recovered point

Execute an isolated etcd restore, interpret the recovered point, and distinguish integrity, identity, and business acceptance.

Start with the timeline

The lab creates a disposable cluster, writes three keys, and saves a snapshot. It then updates one key, deletes another, and acknowledges twenty additional writes. Only then does it stop the source and restore the earlier artifact. This sequence lets you predict the outcome before observing it: the later update is absent and the subsequently deleted key reappears. Record which operations belong on each side of the captured point. In a funds service, a later business acknowledgement remains relevant evidence even when restored data does not contain it. Reconciliation needs an authorized source to explain that difference.

Inventory the artifact and version

With etcd 3.6, the exercise uses etcdctl snapshot save to obtain the artifact through an endpoint and etcdutl snapshot status to inspect revision and key count. Restore also belongs to etcdutl. The script requires all three tools at version 3.6.15 and records its own digest in the report. It also retains the snapshot digest to check that subsequent operations did not alter it. These controls alone neither authenticate the producer nor link a revision to a business time. An operational manifest should add origin, scope, version, and a known recoverable point, with owners responsible for interpreting them.

Exercise rejection before startup

The exercise creates a snapshot copy and changes only one byte of its integrity trailer. Ordinary restore rejects that copy; skip-hash-check is not used and no process starts from the result. The observation does not assert that no partial files were created. A second control tries restoring into a non-empty directory containing a known sentinel: the attempt fails and the sentinel remains unchanged. Analyze the two results separately. One checks inconsistency in the selected artifact; the other checks protection of an occupied target. Neither authorizes deleting existing data or turning rejection into success by bypassing the control.

Rebuild membership from the same point

All three destination members are restored from the same snapshot into new directories, with new local endpoints and their own token. Restore rewrites member and cluster identities. The lab compares those identities and checks that all three members return expected state. It neither combines snapshots from different times nor asks election to reconcile independent histories. When preparing a change, also establish who controls old writers and which endpoints clients actually use. A new identity distinguishes the recovered cluster, but it does not automatically update a scheduler or prevent an application from contacting the old source.

Run and interpret the lab

Save the complete code below as run.py and use a directory containing etcd, etcdctl, and etcdutl 3.6.15. Run python3 run.py --bin-dir /path/to/binaries --output evidence.json. Permission to bind local ports is required. The script accepts neither existing endpoints nor real data directories: it creates temporary resources, uses synthetic keys, and terminates processes it started. The expected report contains ten passed groups and zero failures, including the next lesson’s exercises. If a failure occurs, retain the result and investigate the specific phase. Do not replace a failed observation with a manually assigned success count.

Accept the function and record what remains

The restored cluster receives a new synthetic write, which is read through another member. This result confirms a limited ability to operate after restore. It does not establish real credentials, loaded capacity, business API behavior, or RTO compliance. In pairs, prepare a record containing the recovered point, missing operations, passed controls, dependencies still to validate, and the acceptance owner. One participant represents APS and the other represents business. Reserve ten minutes to discuss whether degraded operation is possible and which effects require reconciliation before the next window. Keep criteria defined before the exercise visible throughout the decision.

"""Original snapshot recovery lab: only its own disposable loopback etcd processes."""
import argparse
import base64
from datetime import datetime, timezone
import hashlib
import json
import os
from pathlib import Path
import socket
import subprocess
import sys
import tempfile
import time
import uuid


def until(action, timeout=15):
 end = time.monotonic + timeout
 last = None
 while time.monotonic < end:
 try:
 result = action
 if result:
 return result
 except (AssertionError, RuntimeError, subprocess.TimeoutExpired) as exc:
 last = str(exc)
 time.sleep(.15)
 raise AssertionError('Condition not observed before deadline: ' + str(last))


class Cluster:
 def __init__(self, folder, binary_dir):
 self.root, self.bin = Path(folder), Path(binary_dir)
 self.token = 'dr-local-' + uuid.uuid4.hex
 self.procs, self.logs, self.clients, self.peers = {}, [], [], []
 reservations = []
 for _ in range(6):
 s = socket.socket; s.bind(('127.0.0.1', 0)); reservations.append(s)
 self.clients = ['http://127.0.0.1:' + str(s.getsockname[1]) for s in reservations[:3]]
 self.peers = ['http://127.0.0.1:' + str(s.getsockname[1]) for s in reservations[3:]]
 for s in reservations:
 s.close
 self.members = ','.join(f'n{i}={self.peers[i]}' for i in range(3))
 # Ignore inherited etcd configuration; never address external endpoints.
 self.env = {k: v for k, v in os.environ.items if not k.startswith(('ETCD_', 'ETCDCTL_'))}

 def start(self, i):
 assert i not in self.procs or self.procs[i].poll is not None
 log = open(self.root/f'n{i}.log', 'ab'); self.logs.append(log)
 args = [str(self.bin/'etcd'), '--name', f'n{i}', '--data-dir', str(self.root/f'n{i}'),
 '--listen-client-urls', self.clients[i], '--advertise-client-urls', self.clients[i],
 '--listen-peer-urls', self.peers[i], '--initial-advertise-peer-urls', self.peers[i],
 '--initial-cluster', self.members, '--initial-cluster-token', self.token,
 '--initial-cluster-state', 'new', '--heartbeat-interval', '100', '--election-timeout', '1000',
 '--log-level', 'error']
 self.procs[i] = subprocess.Popen(args, env=self.env, stdout=log, stderr=log)

 def stop(self, i, abrupt=False):
 p = self.procs[i]
 if p.poll is None:
 p.kill if abrupt else p.terminate
 p.wait(timeout=5)

 def ctl(self, i, *args, raw=False, data=None, expect=True):
 r = subprocess.run([str(self.bin/'etcdctl'), '--endpoints='+self.clients[i], '--dial-timeout=1s',
 '--command-timeout=2s', '--write-out=json', *args], input=data,
 text=True, capture_output=True, timeout=5, env=self.env)
 if expect and r.returncode:
 raise RuntimeError(r.stderr[-600:])
 return r if raw else json.loads(r.stdout)

 def get(self, i, key, serial=False):
 args = ('get', key, '--consistency=s') if serial else ('get', key)
 data = self.ctl(i, *args)
 kv = data.get('kvs', [])
 return None if not kv else base64.b64decode(kv[0]['value']).decode

 def status(self, i):
 return self.ctl(i, 'endpoint', 'status')[0]['Status']

 def leader(self, members):
 def find:
 states = [(i, self.status(i)) for i in members]
 ids = {s.get('leader', 0) for _, s in states}
 if len(ids)!= 1 or 0 in ids:
 return False
 return next(([i] for i, s in states if s['header']['member_id'] == s['leader']), False)
 return until(find)[0]

 def close(self):
 for i in self.procs:
 self.stop(i)
 for log in self.logs:
 log.close


def run(binary_dir):
 binary_dir=Path(binary_dir)
 version=subprocess.run([str(binary_dir/'etcd'),'--version'],text=True,capture_output=True,check=True).stdout
 utility=subprocess.run([str(binary_dir/'etcdutl'),'version'],text=True,capture_output=True,check=True).stdout
 assert 'etcd Version: 3.6.15' in version and 'etcdutl version: 3.6.15' in utility
 checks=[]
 def record(name,**values):checks.append(dict(name=name,passed=True,observations=values))
 with tempfile.TemporaryDirectory(prefix='dr-recovery-etcd-')as folder:
 root=Path(folder);clusters=[]
 def cluster(name):
 p=root/name;p.mkdir;c=Cluster(p,binary_dir);clusters.append(c);return c
 def utl(*args,expect=True):
 result=subprocess.run([str(binary_dir/'etcdutl'),*map(str,args)],text=True,capture_output=True,timeout=20)
 if expect and result.returncode:raise RuntimeError(result.stderr[-1000:])
 return result
 def restore(c,snapshot,bump=None):
 for i in range(3):
 extra=[]if bump is None else ['--bump-revision',str(bump),'--mark-compacted']
 utl('snapshot','restore',snapshot,'--name',f'n{i}','--data-dir',c.root/f'n{i}',
 '--initial-cluster',c.members,'--initial-cluster-token',c.token,
 '--initial-advertise-peer-urls',c.peers[i],*extra)
 for i in range(3):c.start(i)
 until(lambda:all(c.ctl(i,'endpoint','health')[0]['health']for i in range(3)))
 try:
 source=cluster('source')
 for i in range(3):source.start(i)
 until(lambda:all(source.ctl(i,'endpoint','health')[0]['health']for i in range(3)))
 leader=source.leader(range(3))
 source.ctl(leader,'put','/dr/day','2026-10-04')
 source.ctl(leader,'put','/dr/config','before-snapshot')
 source.ctl(leader,'put','/dr/retained','present-at-snapshot')
 snapshot=root/'snapshot.db'source.ctl(leader,'snapshot','save',str(snapshot),raw=True)
 digest=hashlib.sha256(snapshot.read_bytes).hexdigest
 status=json.loads(utl('snapshot','status',snapshot,'--write-out=json').stdout)
 revision=status['revision'];assert revision>=4
 record('live-snapshot-and-status',snapshotCreated=True,revisionObserved=True,atLeastThreeKeys=status['totalKey']>=3,sha256Recorded=True)
 source.ctl(leader,'put','/dr/config','after-snapshot')
 source.ctl(leader,'del','/dr/retained')
 for i in range(20):source.ctl(leader,'put','/dr/after/'+str(i),'acknowledged-after-snapshot')
 source_state=source.status(leader);latest=source_state['header']['revision'];original_cluster=source_state['header']['cluster_id']
 original_members={m['ID']for m in source.ctl(leader,'member','list')['members']}
 assert latest>revision and hashlib.sha256(snapshot.read_bytes).hexdigest==digest
 record('source-advances-after-snapshot',laterRevisionGreater=True,postSnapshotWriteAcknowledged=True,snapshotUnchanged=True)
 for i in range(3):source.stop(i)
 corrupt=root/'corrupt.db'data=bytearray(snapshot.read_bytes);data[-1]^=1;corrupt.write_bytes(data)
 bad=utl('snapshot','restore',corrupt,'--data-dir',root/'rejected-corrupt',expect=False)
 assert bad.returncode!=0 and ('hash' in bad.stderr.loweror'sha256'in bad.stderr.lower),bad.stderr
 record('corrupted-integrity-trailer-rejected',restoreSucceeded=False,hashCheckBypassed=False,processStarted=False)
 occupied=root/'occupied'occupied.mkdir;sentinel=occupied/'keep.txt'sentinel.write_text('owned-fixture-must-survive')
 exists=utl('snapshot','restore',snapshot,'--data-dir',occupied,expect=False)
 assert exists.returncode!=0 and sentinel.read_text=='owned-fixture-must-survive',exists.stderr
 record('existing-nonempty-target-rejected',restoreSucceeded=False,sentinelPreserved=True)
 plain=cluster('plain');restore(plain,snapshot)
 for i in range(3):
 assert plain.get(i,'/dr/config')=='before-snapshot'
 assert plain.get(i,'/dr/retained')=='present-at-snapshot'
 assert plain.get(i,'/dr/after/0')is None
 state=plain.status(0);members=plain.ctl(0,'member','list')['members']
 assert state['header']['cluster_id']!=original_cluster and not original_members.intersection(m['ID']for m in members)
 record('same-snapshot-restores-new-logical-cluster',members=3,originalMemberIdsAbsent=True,newClusterIdentity=True,preSnapshotValueRecovered=True,postSnapshotWriteAbsent=True,postSnapshotDeletionNotReplayed=True)
 assert state['header']['revision']==revision and state['header']['revision']<latest
 future=plain.ctl(0,'get','/dr/config','--rev='+str(latest),raw=True,expect=False)
 assert future.returncode!=0 and 'future revision'in future.stderr.lower,future.stderr
 record('plain-restore-regresses-revision',restoredAtSnapshotRevision=True,belowPreviouslyObservedRevision=True,futureRevisionReadRejected=True)
 plain.ctl(0,'put','/dr/isolated-probe','plain-new-write')
 assert plain.get(1,'/dr/isolated-probe')=='plain-new-write'
 record('ordinary-restored-cluster-accepts-new-operation',syntheticWriteAcknowledged=True,readFromAnotherMember=True)
 for i in range(3):plain.stop(i)
 bumped=cluster('bumped');restore(bumped,snapshot,bump=1000)
 current=bumped.status(0)['header']['revision']
 assert current>=revision+1000 and current>latest
 assert bumped.get(0,'/dr/config')=='before-snapshot' and bumped.get(0,'/dr/after/0')is None
 record('revision-bump-does-not-recover-later-data',revisionAbovePriorObservation=True,bump=1000,preSnapshotValueRecovered=True,postSnapshotWriteAbsent=True)
 historic=bumped.ctl(0,'get','/dr/config','--rev='+str(latest),raw=True,expect=False)
 assert historic.returncode!=0 and 'compacted'in historic.stderr.lower,historic.stderr
 args=[str(binary_dir/'etcdctl'),'--endpoints='+bumped.clients[0],'--write-out=json','watch','/dr/config','--rev='+str(latest)]
 watch=subprocess.Popen(args,text=True,stdout=subprocess.PIPE,stderr=subprocess.PIPE,env=bumped.env)
 try:
 try:out,err=watch.communicate(timeout=4)
 except subprocess.TimeoutExpired:
 watch.terminate;out,err=watch.communicate(timeout=3)
 assert 'compacted'in(out+err).lower,(out,err)
 finally:
 if watch.pollis None:watch.kill;watch.wait(timeout=3)
 record('compacted-history-signals-consumer-reset',historicReadRejected=True,oldRevisionWatchReportedCompaction=True,realKubernetesInformerExecuted=False)
 write=bumped.ctl(0,'put','/dr/isolated-probe','bumped-new-write')
 assert write['header']['revision']>current
 until(lambda:all(bumped.get(i,'/dr/isolated-probe')=='bumped-new-write'for i in range(3)))
 assert hashlib.sha256(snapshot.read_bytes).hexdigest==digest
 record('bumped-cluster-new-write-and-snapshot-preserved',newWriteRevisionIncreased=True,threeMemberReadAgreement=True,originalArtifactUnchanged=True)
 finally:
 for c in clusters:c.close
 return dict(executedAt=datetime.now(timezone.utc).isoformat,version=version.strip,utilityVersion=utility.strip,passed=len(checks),failed=0,checks=checks,scriptSha256=hashlib.sha256(Path(__file__).read_bytes).hexdigest,scope='Original local snapshot save, integrity rejection, ordinary three-member restore, revision-bumped restore and compacted-watch observations with etcd/etcdctl/etcdutl 3.6.15. Sequential disposable clusters on one Darwin ARM64 host, loopback HTTP, synthetic keys. No existing cluster, real disaster, Kubernetes informer, physical fencing, production workload, credential/key recovery, business API or RPO/RTO benchmark.')


if __name__=='__main__':
 p=argparse.ArgumentParser;p.add_argument('--bin-dir',required=True,type=Path);p.add_argument('--output',type=Path);a=p.parse_args;result=run(a.bin_dir);payload=json.dumps(result,indent=2)+'\n'
 if a.output:a.output.write_text(payload)
 print(payload)
IN PRACTICE

A rule is captured at 18:00 and deleted at 18:04. Restoring the 18:00 point brings it back. The team uses the authorized change record to decide reconciliation without treating reappearance as a new business instruction.

Common pitfalls

Confusing a hash with authenticated provenance; using different snapshots per member; ignoring subsequent deletions; accepting a synthetic write as validation of the entire service.

Related topics: Objectives and dependencies · Restore: recovered point and integrity · Resumption: replay and operational acceptance

Take this idea with you

A restore should demonstrate the artifact used, the point produced, the identity created, and the function validated. Replica agreement does not recover information after the snapshot.

Create account

Reference: Disaster recovery · BigSavant recovery 2026-09; PostgreSQL 18, etcd 3.6 and selected AWS/Azure behavior