Concept and mechanism
A service can maintain its SLA after losing a path while still having degraded redundancy. The alert should preserve that distinction: current access does not establish margin for the next failure. Correlate path states, host and fabric events, LUN identity, and application symptoms. When every path disappears, a queue policy can appear as a hung batch rather than an immediate error. Indiscriminately retrying or restarting does not guarantee recovery. Decide with owners how to contain impact, restore access, and reconcile functional outcome. Switching between waiting and returning errors has consequences that belong in the runbook.
Guided application
Persistent reservations restrict access according to type and cluster context. In the interface documented by the Linux kernel, PR_WRITE_EXCLUSIVE allows other initiators to read while reserving writes to the holder; other types have different rules. A reservation conflict should not be blindly removed. In a fictional partial-failure scenario, the secondary does not know whether the primary still writes. Before taking authority, confirm exclusion of the previous writer through the supported mechanism. A peer unreachable over the network does not establish that it lost SAN access. Preserve evidence and avoid clearing every key as a universal response: that can remove the control preventing improper concurrency.
An available service with a faulty path needs redundancy restored even without a user complaint.
Common pitfalls
Current SLA as full resilience; unreachable peer as stopped writer; reservation conflict as a purposeless error.
Related topics: Topology and identity · Fibre Channel and LUN access · iSCSI: sessions and security
Recover the path and write authority with evidence of functional state.
Reference: Linux block-layer persistent reservations · DR SAN 2026-09; selected RHEL 9, ONTAP 9 and iSCSI behavior