Start with service outcome
A running process and a 200 endpoint do not confirm that the current cycle’s report is available. Identify user population, operation, required freshness, and expected outcome. If samples and collector heartbeat are both missing, record an observability gap instead of concluding zero errors. Seek alternative sources and investigate collection. During a release, separate version and operation type: the global average can improve while writes in the new cohort fail. APS work connects those signals to impact and the next decision while preserving each observation’s scope and limitations.
Organize facts, hypotheses, and chronology
A change immediately preceding an incident is a clue, not automatically a confirmed cause. Record the observation and formulate a hypothesis distinguishable from alternatives. Retain timestamps with offsets: 10:55+01:00 is 09:55 UTC, so at 10:00 UTC the result is five minutes old. The laboratory calculates that difference using timezone-aware instants; it does not validate clock synchronization. In L3 escalation, include impact, window, population, relevant changes, attempted actions, and outcomes. A restart reducing errors can provide useful mitigation, but does not automatically reconcile earlier operations or conclude causal investigation.
Hand over responsibility for current state
The handover model stores the incoming team, accepted revision, and current revision. Team B accepted revision 4; revision 5 then added unknown operations and a replay restriction. Comparison marks handover incomplete. This is a simple model exposing a gap, not an actual approval application. In practice, explain the material change, confirm acceptance, and communicate who coordinates. Include reconciliation owners, next steps, and time of the business update. If the receiver is unavailable, address continuity of coverage through the local process. Shift end or sending a link does not by itself demonstrate effective transfer.
Demonstrate autonomy in the intended role
A supplier recovered the service using its own account. The intended RUN operator cannot access the required resource and has not accepted the service. The readiness model keeps ready false even with recoveryTested true. The lesson is to assess capability in the context of those becoming responsible. A runbook should specify prerequisites, scope, expected outcomes, checks, and stop or escalation limits. Request controlled execution by the authorized role without relying on the author to interpret every step. If transitional support exists, define coverage, duration, limits, and accountable acceptance. The laboratory tested no actual permissions and conducted no human acceptance session.
Forecast recovery within cutoff
The model has 1200 accumulated items, arrivals of 30 per minute, and service capacity of 70 per minute. The queue shrinks at 40 per minute, taking 30 minutes to drain. Cutoff is 25 minutes away and final validation requires five, leaving only twenty for processing. Required rate would be 90 per minute: sixty net reduction plus thirty arrivals. This is a calculated requirement, not demonstrated capacity. Evaluate admission, priority, capacity, or permitted partial-result options with accountable owners. Adding workers does not guarantee that gain when another dependency limits service. Constant rates are an explicit exercise assumption.
Connect support, project, and improvement
The 75% support and 25% project split belongs to the role described in this context; it is not a universal APS or SRE rule. If incidents consume planned project capacity, expose the deviation and review commitments. Do not retain dates based on a capacity assumption that is no longer true or reclassify hours to hide the problem. Close the workshop with an English summary: known impact, uncertainties, mitigation, pending decisions, owner, and next update. Add an operational gap to the improvement plan with a verifiable completion criterion. The guide is ready for team practice, but the human session and specialist review remain outstanding.
backlog = 1200
arrivals_per_minute = 30
service_per_minute = 70
drain_minutes = backlog / (service_per_minute - arrivals_per_minute) # 30
cutoff_minutes = 25
validation_minutes = 5
required_service = arrivals_per_minute + backlog / (cutoff_minutes - validation_minutes) # 90/min
# Synthetic constant-rate model, not measured capacity or an approved action.The next team accepted incident revision 4; revision 5 adds unknown operations and a replay restriction. Earlier acceptance does not cover that new state.
Common pitfalls
Confusing missing metrics with zero errors, global average with every cohort being healthy, document delivery with acceptance, or supplier demonstration with RUN autonomy.
Related topics: Incident Management · Technical Project Management · Monitoring and Observability
An autonomous team knows current state, can execute authorized procedures, and explicitly accepts outstanding gaps and next decisions.
Reference: Managing Incidents · BigSavant APS professional curriculum 2026-09; vendor-neutral operational guidance reviewed 2026-09-30