Start with impact and the required decision
At 05:00, a batch fails and financial close is at 06:00. Gather facts about affected work, consumers, deadline, and available alternatives. Alert count or the caller’s seniority does not replace agreed impact and urgency criteria. The exercise supplies a local matrix: blocked payments within 30 minutes of cut-off are P1; reports with a workaround until the next day are P2. This rule serves the fictional case, rather than being a universal ITIL table. The first decision should reduce impact within authorized scope.
An active process still needs acceptance
Define what counts as recovery before confusing a technical indicator with the service. For this batch, criteria are a complete file, no duplicates, and confirmed consumption. A restart may be useful without satisfying all three. Keep every check’s status visible. If some transactions were already applied, replay can create duplicate effects. The team needs to reconcile outcomes and choose a supported recovery; repeatedly rerunning until the scheduler error disappears is insufficient.
Restoration and investigation can proceed in parallel
A valid workaround can reduce impact while problem management continues investigating actual or potential causes. Link recurring incidents and retain conditions under which the alternative was rehearsed. Version, configuration, external effects, and authorization can limit reuse. Do not declare the cause removed merely because one run completed. Also avoid withholding viable recovery solely to await a complete explanation. The record should let the next shift understand what works, what remains uncertain, and residual risk.
Preauthorization has boundaries
The title certificate renewal does not by itself determine change handling. In this exercise, the standard model covers version A and the same algorithm. A proposal using B and changing algorithm exceeds the supplied conditions. Assess risk, testing, recovery, and appropriate authority. Monthly frequency does not automatically expand the model. If imminent expiry makes intervention urgent, the local procedure allows contact with an emergency authority; that expedites the decision without turning urgency into implicit approval.
Communicate without inventing a forecast
A useful update distinguishes impact, confirmed facts, hypothesis, and next decision. At 05:20, you can report resumed processing and pending reconciliation, with an update at 05:30. That communication commitment is not a promise of recovery by 05:30. If the team lacks a reliable estimate, state what is needed to produce one. The service desk maintains contact and understanding of the need while specialists analyze the mechanism. Avoid making the customer reconstruct the incident from multiple supplier tickets.
Shift handover transfers the ability to act
A handover should enable the colleague to make the next decision. Include scope, chronology, already-produced effects, necessary access, available authority, and recovery criteria. Check out-of-hours coverage: receiving an alert is insufficient if the only supplier able to recover the database responds only during office hours. The gap spans people, information, partner coverage, and flow. Documentation can be complete while operational capability remains absent. Identify an owner and resolve that difference before accepting 24/7 operation.
Guided practice: three records, three purposes
In the fixture, the incident tracks functional restoration, the problem tracks recurrence and cause, and the change records the intervention decision. These are different purposes, rather than three mandatory teams or a rigid sequence. Identify what may close under local criteria and what remains open. Guiding answer: running state and external ticket closure do not satisfy recovery; the cause also remains unresolved. Complete the matrix with an owner and necessary evidence. The exercise does not execute an ITSM workflow or know any bank’s internal procedures.
Synthetic service fixture, not an ITSM workflow
05:00 incident: batch stopped; cut-off 06:00
05:20 process=running; reconciliation=pending
consumer: some transactions already applied
problem: cause unresolved; recurrence observed
change: restart authorized; replay not yet assessed
next communication: 05:30, not a recovery promiseThe job resumed at 05:20, but reconciliation is pending and partial application occurred. The owner communicates status and coordinates validation before approving replay or full recovery.
Common pitfalls
Closing based on running state; replaying uncertain effects; applying an out-of-version workaround; confusing emergency with authorization; handing over a runbook without access or coverage.
Related topics: Incident management · Service levels and improvement
Recovery, cause, and change need their own criteria and sufficient information for the next decision.
Reference: Incident Management practice overview · ITIL 4 Foundation; syllabus v4.2.0 (March 2025)