A queue persists after a path returns
The fifth group starts with 300 waiting operations. Arrivals are 80 per second and completions are 100 per second, without new failures, additional limits, or retries. Net capacity for reducing the queue is twenty, so the model predicts fifteen seconds. After five seconds, two hundred operations remain. Calculate the balance including arrivals and departures before consulting the result. Dividing 300 by 100 ignores new work and produces an overly optimistic forecast. These numbers were chosen to teach reasoning; they are not SAN measurements. During an incident, define a comparable window, collect actual rates, and revise the forecast when workload or capacity changes.
Distinguish an unknown deadline from an empty queue
The next variant retains 300 operations but makes arrivals and capacity equal at 80 per second. The drainSeconds=null result means the model finds no net reduction, not that the queue is empty. A positive completion rate can coexist with persistent delay. In reporting, provide observed backlog, the difference between rates, and forecast conditions. Avoid promising a deadline when available capacity does not exceed demand. With the owners, evaluate measures such as admission control or increased effective capacity, considering business impact. The exercise does not automatically select a mitigation. That choice requires priorities, limits, and evidence absent from the simple arithmetic.
Reconcile unacknowledged writes
A path returning does not, by itself, clarify each operation’s outcome. In the fictional case, three writes still lack application confirmation even after a successful probe. Do not automatically classify them as unexecuted or repeat them with new identifiers without understanding consumer semantics. Record available business identifiers, the affected interval, and who confirms results. An I/O queuing policy and an application timeout can create different perspectives on the same work. The workshop uses that distinction to practice communication and reconciliation. It does not simulate the write protocol, database durability, or guarantees of a payment processor.
Verify expansion by layer
The sixth group defines an intended size of 750 GiB and four observations: 750, 750, 500, and 500. Its consistency predicate fails. When all four values become 750, that predicate passes, but the model still states that filesystem growth was not established. Each piece of evidence has limited scope. During real expansion, identify the effective stack, including any partitions, LVM, encryption, and consumers, and follow the applicable procedure. Do not add path capacities or use their average as an acceptance criterion. The script performs no rescan and resizes no devices. The workshop product is a layer-by-layer verification list with owners and conditions for stopping the change.
Build a valid withdrawal order
The seventh group uses an original task graph. Stopping API and batch are prerequisites for releasing consumers; flush-and-detach and mapping withdrawal follow. The ordering respects both inputs. A variant withdraws mapping immediately after stopping the API and is rejected because batch and subsequent layers remain unhandled. Another variant contains a cycle and admits no complete ordering. The result helps review a plan but does not establish that any task was executed. Before decommissioning, expand each broad step into a platform-appropriate procedure and define completion evidence. An old schedule can remain a consumer even when no process is active during that minute.
Close with acceptance and explicit outstanding work
In the last group, all four expected paths are present and the I/O probe passed. Batch acceptance is false and three writes await reconciliation, so the fictional closure criterion remains false. Do not discard evidence of technical recovery, but do not declare complete functional recovery either. Finish the workshop with a five-line report: current impact, collected evidence, outstanding work, owner, and next update. Then review a scenario where completion rate falls during recovery and explain why the previous forecast no longer applies. The guide is prepared for a team discussion; this publication does not claim that a human group executed or validated it.
python3 content/labs/san-evidence/run.py
# Synthetic queue: initial=300, arrivals=80/s, completions=100/s
# Net drain=20/s; after 5s: 200 pending; total drain=15s
# If arrivals=capacity=80/s: no finite drain forecast
# Growth fixture: [750,750,500,500] GiB -> path gate fails
# Offline calculations; no multipath configuration or LUN changes.After restoring paths for a fictional financial-instruction application, queued work and unacknowledged writes remain. Reporting distinguishes connectivity from functional acceptance.
Common pitfalls
Dividing backlog by gross capacity, promising times without assumptions, treating timeout as an unexecuted write, and withdrawing mapping before stopping consumers.
Related topics: Incident Management · Change Management · Disaster Recovery
Recovery can be described accurately only when paths, pending work, unknown outcomes, and service acceptance are distinguished.
Reference: fractions: rational numbers · BigSavant SAN 2026-09; selected RHEL 9, ONTAP 9 and iSCSI behavior