Start with each read requirement
A fictional application confirms fund subscriptions and produces reports. Confirmation must observe the commit; reports accept up to 30 seconds of lag. These accesses should not automatically receive identical routing. An asynchronous PostgreSQL read replica can scale reports, but an immediate read may still return the previous value. Define primary confirmation transaction semantics and a report policy when lag exceeds the limit. Replica availability does not establish sufficient freshness for every consumer.
Identify the Multi-AZ topology
Multi-AZ does not describe a single topology. In an RDS Multi-AZ DB instance deployment with one standby, that standby does not serve application reads. A Multi-AZ DB cluster has two standbys that can serve reads. Before recommending a reporting target, identify deployment type, engine, and regional availability. Adding a DNS alias does not change a standby’s role. During architecture review, draw which components accept writes, reads, and failover instead of using only a box labelled Multi-AZ.
Interpret the Aurora reader endpoint
The Aurora reader endpoint distributes connections, not each SELECT within a session. A client with one persistent connection can concentrate a thousand queries on one replica while another remains lightly used. Examine pooling and session duration before declaring a balancing failure. Connection distribution also does not guarantee equal CPU because queries have different costs. Without Aurora Replicas, the reader endpoint connects to the primary and can allow writes; its name does not replace SQL permissions and access design.
Use RDS Proxy for the right problem
A burst of functions with short connections can benefit from connection reuse through RDS Proxy. This does not turn the proxy into a result cache or fix a query scanning millions of rows. With few connections and query-dominated CPU, examine the plan and indexes. Session state can cause pinning and limit reuse; behavior depends on engine and operation. Measure connections, pinned sessions, latency, and database load before attributing extra capacity to the proxy.
Calculate an explicit concurrency bound
In a fictional model, the database allows 240 sessions, 40 are reserved, and each worker occupies two without multiplexing. The remaining 200 sessions allow at most 100 workers by this resource limit. This is not a throughput promise: CPU, locks, memory, and latency can impose a lower bound. If session state prevents reuse, do not assume the proxy reduces the calculation. Record assumptions in the capacity plan and measure representative load before production approval.
Restore data and recover service
RDS point-in-time restore creates a new instance; it does not automatically rewrite the old instance or change application connections. After a deletion at 10:05, restoring to 10:04 may recover lost data but also excludes legitimate later operations. Preserve evidence, constrain concurrent writes according to the plan, and reconcile those operations. Validate data, permissions, and functional outcomes before controlling connection cutover. Available status is a technical condition rather than complete evidence of business recovery.
240 sessions minus a 40-session reserve, divided by two sessions per worker, gives 100 workers in the model.
Common pitfalls
Confusing standby with a read replica, connection balancing with query balancing, or a restored instance with a recovered application.
Related topics: Availability and recovery · Database performance
Choose each mechanism for the requirement it addresses and validate the application after change.
Reference: RDS Multi-AZ deployment types · SAA-C03