← Disaster Recovery: prepare, recover, and validate
07 / 8 · 60 MIN

Restore: recovered point and integrity

Distinguish a readable file from a recovered dataset using backups and counterexamples executed locally.

Define the reference before restoring

The lab uses a fictional application with two orders, for 100 and 200 units, on business day 2026-10-02. Each order has lines whose total must match its amount. These values form the exercise reference and are defined outside the copy being assessed. If the reference were calculated only from restored data, a lost order would also disappear from the expected result. In APS work, identify the available independent control: accepted identifiers, input totals, or confirmation from the preceding process. Document scope, business date, and reference origin before interpreting a green report.

An incomplete copy may remain readable

In the first exercise, the initial order is committed and the lab performs a checkpoint. It then commits the second order while retaining the open connection and WAL writes. It deliberately copies only the main file to demonstrate an unsuitable method under these conditions. The source has two orders, but that copy contains one; integrity_check returns ok. A copy obtained through the backup API contains both. The result is not a recommendation to copy live files manually. It is a controlled observation of a check passing against an earlier point. Record artifact integrity separately from coverage of the commitments that recovery was intended to preserve.

Separate ongoing work from committed records

The second exercise opens a transaction and inserts a third order. The writer sees three while a backup made through a separate connection contains two. The original transaction ends with rollback. This difference is expected: the third order was never committed for that reader. The backup connection is not the one holding the uncommitted write. This distinction matters when someone compares an open session screen with a recovered artifact. Before concluding data loss, establish whether the operation committed, which point was observed, and which evidence belongs to the transaction. The lab does not simulate power failure or measure physical storage guarantees.

Control artifact, scope, and destination

The local guard compares a closed backup hash with a trusted expected value. A truncated file is rejected before creating the destination. Another exercise uses valid bytes but a different scope: exercise-a cannot be accepted as exercise-b. A third path finds an existing destination and preserves it unchanged. These are three distinct questions: did we receive the expected artifact, does it belong to the intended scope, and may we use this destination? A digest does not independently authenticate a file if an attacker can also replace the expected value. The code assumes a private temporary directory, stable files, and no concurrent path changes; it is not a restore tool for hostile systems.

Check references and business rules

The lab creates two defects in disposable copies. In the first, it inserts a line with a missing reference, deliberately disabling enforcement only to construct the counterexample. integrity_check passes and foreign_key_check finds one violation. In the second, references are valid but lines for the 200-unit order total 190. Both technical checks pass and amount reconciliation fails. Acceptance requires checks that match the data contract. Do not adjust the value merely to pass the test: retain the discrepancy, seek the authorized source, and retest an approved correction. A readable database can still be unsuitable for producing the closing file.

Reproduce and limit the conclusion

The run.py file creates every database in a temporary directory and accepts no existing database paths. Published evidence records Python 3.13.1 and SQLite 3.53.4. The SQLite version used is later than the WAL fix described in its documentation; the exercise does not attempt to reproduce that race condition. Eight groups check local operations, and the script hash ties evidence to executed code. No result establishes PostgreSQL, Oracle, cloud recovery, or real failover. To transfer learning into a project, define mechanisms for the actual technology and repeat equivalent checks in an authorized environment. Retain versions, expected data, results, and limits with the acceptance report.

python3 content/labs/recovery-restore/run.py
# Requires SQLite >= 3.51.3; recorded run: Python 3.13.1 / SQLite 3.53.4
# All databases are disposable local fixtures.
IN PRACTICE

The source has two orders. A main-file-only copy has one and passes integrity_check; the API copy has two. The recovered point must be checked separately.

Common pitfalls

Calculating the reference only from the restore; copying only the main file in WAL; treating a digest as authorization; correcting amounts without a business source.

Related topics: Resumption: replay and operational acceptance · RTO, RPO, and dependencies · Exercises and functional validation

Take this idea with you

Accept a restore by its recovery point and data contract, using complementary checks and reproducible evidence.

Create account

Reference: SQLite Backup API · BigSavant recovery 2026-09; PostgreSQL 18, etcd 3.6 and selected AWS/Azure behavior