Concept and mechanism
Transition begins during design, while operational requirements can still be corrected with less rework. Involve RUN, development, infrastructure, and security in defining readiness. Review observability, access, procedures, dependencies, capacity, recovery, escalation, and training. Delivered documentation does not prove operational competence: exercises and demonstrations help establish autonomy. The Production Readiness Review model described by Google SRE is a contextual reference, not a universal requirement. Define which gaps prevent transition, which can receive explicit acceptance, and who owns risks and actions. Operational responsibility should transfer in a way acknowledged by the teams involved.
Guided application
In a fictional weekend cutover, prepare sequence, duration, owners, contacts, go/no-go criteria, checkpoints, and recovery. A reserved window does not replace readiness. If a data change is incompatible with the previous version, reverting only the binary can fail: rehearse the data and reconciliation strategy. During progressive exposure, use metrics with representative traffic; absence of errors without usage is weak evidence. Define post-launch support, exit criteria, and support acceptance. The project should make operations sustainable, including batch work and external consumers that may reveal problems later.
Go/no-go needs evidence, authority, and a workable recovery path.
Common pitfalls
Window as authorization; code-only rollback; delivered runbook as demonstrated autonomy.
Related topics: Mandate, scope, and acceptance · Capacity, dependencies, and forecasting · Risk, change, and architecture decisions
Validate the complete workflow and the team that will operate it.
Reference: The evolving SRE engagement model · PM² reference practices and cloud operational governance; primary guidance reviewed 2026-09-30