Accept operations as well as documentation
A team may have received the manual and still be unable to operate the service. Define representative tasks demonstrating access, diagnosis, action, and escalation. In a fictional exercise, the night shift must identify the affected batch, inspect authorized logs, and locate the reconciliation owner without relying on a project member’s personal account. Record who executed each task, in which environment, and with what result. One person’s successful exercise does not establish coverage across all shifts. Findings need correction, ownership, and rechecking; responsibility transfer uses only evidence that actually exists.
Measure the outcome that matters
Define the indicator through its population, good outcome, time window, and data source. A service returning HTTP 200 but delivering an incomplete file can miss the business objective. If 98 of 100 eligible files arrived complete before cutoff, that exercise’s rate is 98%. Excluding the two late files from the denominator changes the measured question. Distinguish indicator, target, and service agreement rather than inventing contractual penalties. If measurement observes only requests reaching the backend, identify what may be missed before that point and agree sufficient coverage with the owners.
Test alerting, routing, and response
A monitoring rule does not prove that someone receives and handles its alert. In an authorized exercise, check trigger condition, recipient, coverage hours, acknowledgement, and escalation. A green dashboard can coexist with a disabled paging channel. Also verify that the runbook links symptoms to the service and defines operator authority. An alert without an owner creates hidden dependence on the project. If external support operates only during business hours, represent that limitation in promised coverage and decide how to respond outside those hours before accepting continuous availability.
Rehearse recovery through usable service
A completed backup does not demonstrate usable restoration. In a rehearsal, verify destination access, data, required keys, application, dependencies, and functional validation. Measure from the agreed starting point to the accepted outcome; do not stop timing at machine startup if reconciliation is included. A tabletop exercises discussion and coordination but does not by itself establish data-copy duration. Identify common components that can fail across both locations, such as identity or key management. Record tested scope, results, and exclusions rather than turning a partial demonstration into a resilience guarantee.
Manage exceptions and end temporary support
An accepted exception needs scope, risk, temporary control, owner, deadline, and closure evidence. Do not use an issue list as a substitute for acceptance by the operating owner. If the project remains in temporary support, define continuing tasks and exit criteria. In an exercise, exit may require two full cycles without project intervention and a successful escalation rehearsal; those are case-specific criteria rather than a universal standard. If unmet, propose an authorized extension or correction. Reaching the planned date does not automatically make the team autonomous.
Maintain the service after project closure
Delivery continues changing after project closure. Assign owners to review contacts, access, monitoring, recovery procedures, and dependencies after relevant changes. In a fictional service, replacing middleware may invalidate runbook steps and previous rehearsal timings; the record should trigger review proportionate to impact. Link obsolescence and decommissioning work to teams that accepted responsibility and funding. Retain closure evidence, benefits, and limitations. Administrative closure is a governance decision rather than removal of tasks still needed to maintain the service.
Backup testing is green, but RUN cannot obtain the required key at the recovery location. The project has not demonstrated usable restoration. The finding needs authorized access, rehearsal, and acceptance; sending the manual does not resolve it.
Common pitfalls
Treating training attendance as demonstrated competence; measuring components only; excluding failures from the denominator; confusing tabletop with restoration; ending temporary support solely by date.
Related topics: Lead decisions and communicate what matters · Address risk, obsolescence, and recovery · Prepare the change and hand over autonomy
RUN autonomy is demonstrated through people, access, procedures, and outcomes; it needs maintenance after the project.
Reference: Contingency Planning Guide for Federal Information Systems · DR Technical Project Manager 2026.4; independent professional curriculum