← Release Manager: versions, readiness, and operations
10 / 10 · 60 MIN

Stabilization, handover, and evidence-based closure

Distinguish the end of enhanced support, operational autonomy, residual actions, and demonstrated learning.

Exit enhanced support with explicit conditions

Enhanced support, often called hypercare, brings development and operations together during stabilization. Its calendar end does not establish transferred capability. In a fictional case, APS can diagnose but still cannot recover processing using its own account. Identify the missing task, required access or knowledge, and who responds until autonomy is demonstrated or other coverage is explicitly accepted. Google's public account of bringing services into SRE describes practical learning and progressive responsibility transfer in a specific organizational context. Apply the evidence principle to local responsibilities without imposing Google's model on a bank. Handover may retain accepted defects if impact, mitigation, owner, deadline, and review are clear. “Accepted and open” does not mean “fixed.” An overdue review date needs a current decision; it does not automatically renew earlier acceptance. A calm dashboard is also insufficient if the required recovery task has never been performed by the receiving team.

Communicate states without merging milestones

Distinguish published, installed, activated, and validated states. If the package is installed but its flag remains off and business validation starts only at 09:00, communicate those conditions separately. During recovery, the API may be available at 20:10, the queue complete at 20:45, and reconciliation accepted at 21:00. There are 35 minutes between API availability and queue completion, then fifteen more until reconciliation, but these facts do not establish fifty minutes of total unavailability. Select the measure matching the service and commitment being reported. An aggregate green state that hides pending reconciliation may lead the business to make an incorrect decision. Also check infrequent consumers: three days without traffic do not establish that a monthly consumer no longer needs the old endpoint. Connect decommissioning to dependency evidence, an explicit decision, and applicable retention conditions. Assign an owner to resolve unknown consumers instead of interpreting silence as acceptance.

Turn an incident into a verifiable change

In an incident review, separate facts, causal hypotheses, and actions. An error increase four minutes after release deserves investigation, but a network change also occurred; temporal order alone does not establish cause. Reconstruct the sequence, compare context, and identify what remains unobserved. An action to “improve testing” has no verifiable completion condition. Instead state: “the test owner adds the affected configuration combination and demonstrates error detection before promotion by the agreed date.” A closed ticket without executing that test establishes only administrative status. Learning requires checking the intended effect and reviewing results. Google SRE guidance connects useful reviews with concrete actions that are followed through. Use original examples and select changes tied to contributing conditions; adding generic approval does not itself resolve missing permissions or stale evidence. Preserve the initial observation, the implemented change, and the effectiveness result as separate entries, so later readers can understand the remaining uncertainty.

Close responsibilities and exercise the control again

Two recovery attempts failed because permissions were missing. The team updated the document and closed the ticket, but APS has not yet performed the task. The next useful result is observed execution by the operational account in the applicable context, with failure handling if the issue persists. If other services share the pattern, identify owners and assess applicability before changing them all. Keep residual actions visible after handover, including overdue review dates. A new out-of-scope request should not silently consume capacity reserved for a critical batch: expose priority, impact, and the allocation decision. For the exercise, write an English note stating the observed condition, gap, owner, next evidence, and required decision. Discuss an acceptable alternative, such as extending defined temporary coverage, and its cost. Do not declare every issue resolved merely to close the release. The preceding lesson's Python model demonstrates milestone differences and overdue review, but cannot replace observing these tasks or obtaining the responsible people's acceptance.

Observed state | Limitation | Owner | Next evidence | Deadline/decision
API available 20:10 | backlog pending | APS owner | processing ends 20:45 | update business
Ticket closed | effectiveness untested | test owner | exercise affected combination | agreed date
IN PRACTICE

APS can diagnose but cannot recover; the monthly consumer is still unconfirmed; a residual action is past review. The report retains three rows with owners and decisions instead of a calendar-driven green handover.

Common pitfalls

Treating a closed ticket as demonstrated effectiveness; confusing acceptance with correction; decommissioning from a short traffic sample; attributing cause solely from timing; hiding capacity consumed outside scope.

Related topics: APS handover · Post-incident improvement · Obsolescence and decommissioning

Take this idea with you

Close the release with clear states and responsibilities; check action outcomes before declaring that learning has been applied.

Create account

Reference: The Evolving SRE Engagement Model · Google SRE release and canary guidance; GitHub immutable releases and GitLab release evidence and deployment safety; DORA five-metric model; inspected 2026-10-01