Concept and mechanism
A redundant copy does not resolve every failure mode. In an RDS Multi-AZ DB instance deployment with one synchronous standby, that standby supports availability and does not serve read queries. Do not generalize this behavior to every cluster deployment. Point-in-time restore creates a new instance and requires reviewing connection, parameters, groups, and access. In S3, a live rule covers the applicable flow after creation; earlier objects need a backfill strategy such as Batch Replication. The plan should explain what will be recovered, from which point, and how results will be reconciled before returning to consumers.
Guided application
In Auto Scaling, confirm which checks drive replacement. An unhealthy load-balancer result does not automatically enable ELB integration on the group. Initialization also matters: the grace period avoids certain premature replacements, but does not suspend load-balancer checks or protect an instance that stops running. Measure potential data loss separately from time to usable service. Targets belong to the defined service rather than merely a database available state. For an infrastructure project, include functional testing, client routing, acceptance ownership, and reconciliation evidence in the drill plan.
Disruption at02:00, data through01:52, and service at02:38: eight minutes of potential loss and38 of recovery. With RPO10 and RTO30, only the first target is met.
Common pitfalls
Treating standby as a read replica; counting only technical restoration; assuming live replication copies all history.
Related topics: Signals, alarms, and operational diagnosis · Changes, drift, and controlled automation · Authorization, keys, and secret consumers
Recovery ends when the defined service is usable and data is accepted.
Reference: RDS point-in-time restore · SOA-C02 archived guide v2.3; retired2025-09-29