Identify topology before promising availability
Start by recording engine, version, Region, and deployment mode. In an RDS Multi-AZ DB instance with one standby, that standby supports failover and does not serve reporting queries. In an RDS Multi-AZ DB cluster, one writer and two readers occupy three zones in the same Region; the readers can serve queries. Semisynchronous acknowledgment does not mean both have applied every event. Rehearsal must therefore include reader lag and load. A traditional read replica uses asynchronous replication and can offload reads, but does not establish immediate visibility after a write. In the fictional example, reporting tolerates ninety seconds of staleness; an order acknowledgment cannot tolerate the accepted write being absent. Treat these as separate requirements and measure each through its proposed read path.
Follow endpoint role and renew connections
A cluster writer endpoint follows the current writer; an instance endpoint identifies one particular instance. Using the latter for convenience can leave an application attached to the former writer after a role change. The reader endpoint distributes connection requests, not each query within a persistent session. Measure actual pool distribution before concluding that two readers split work equally. During Multi-AZ DB instance failover, DNS is updated and existing connections need reestablishment. Check JVM DNS caching and pool renewal without fixing IP addresses. In a WebSphere rehearsal, compare resolution for a new connection with addresses used by old sessions. A network test in a shell does not establish the behavior of the JVM running the application.
Use a proxy with an explicit transaction boundary
RDS Proxy can reuse connections when transactions finish and reduce session-establishment work. The benefit depends on traffic actually using the proxy endpoint. Sessions whose state prevents reuse can become pinned; inspect pinned-connection metrics and reasons before indiscriminately increasing pools. Do not remove state required for functional correctness merely to improve a metric. During failover, connections with transactions or statements in progress can be canceled. A proxy does not decide whether to repeat a business order whose outcome is unknown. The runbook should distinguish a repeatable read, an explicitly aborted transaction, and acknowledgment lost after sending COMMIT. In the last case, query a durable identity or reconcile before creating another effect. Retain application errors and outcomes as evidence, not only proxy state.
Measure recovery and rebuild protection
Define measurement start and finish with business stakeholders. In the exercise, service stops responding at 02:10:00, the database accepts new connections at 02:11:20, and the functional flow passes at 02:14:50. That is eighty seconds for the database and 290 seconds for the service; a four-minute objective failed by fifty seconds. These are invented exercise values, not AWS commitments. For planned read-replica promotion, pause writes in a controlled way and demonstrate change application before promotion. After promotion the instance is independent: do not assume previous replication continues or that returning to the old endpoint reverses new writes. The plan includes protection of the new source, write authority, consumer checks, and APS acceptance. Typical documented times help prepare rehearsals; they do not replace measured outcomes.
service_stop = 2 * 3600 + 10 * 60
database_ready = 2 * 3600 + 11 * 60 + 20
service_ready = 2 * 3600 + 14 * 60 + 50
database_seconds = database_ready - service_stop # 80
service_seconds = service_ready - service_stop # 290
objective_seconds = 4 * 60
overrun_seconds = service_seconds - objective_seconds # 50In the fictional rehearsal, Dynatrace shows JDBC errors after the database changes zones. A new connection resolves the correct writer, but the WebSphere pool retains invalid sessions. The team gathers RDS events, JVM resolution, and functional-flow results before adjusting and repeating rehearsal.
Common pitfalls
Confusing the two Multi-AZ modes; treating a reader endpoint as per-query balancing; promising proxy transaction replay; measuring database time alone.
Related topics: Historical restore and recovery evidence
Recovery needs suitable topology, correct reconnection, handling of uncertain outcomes, and functional-flow evidence.
Reference: RDS Multi-AZ DB instance deployments · SAP-C02