← SAN: paths, access, and operations
07 / 8 · 60 MIN

Identity, paths, and access

Correlate volumes across hosts, identify shared failures, and distinguish session, mapping, and reservation in an access decision.

Start with an evidence matrix

In this workshop, inventory rows are fictional and explicitly identified as such. They were not collected from a SAN and do not represent a bank’s infrastructure. Run the local script to obtain eight deterministic results, first comparing predictions with the JSON. For each row, record the host, local alias, WWID, initiator, and traversed components. Add origin and time when working with real evidence. An old inventory can accurately describe a configuration that has since changed. The intended workshop outcome is a reasoned decision accompanied by outstanding checks. Running the model does not establish driver compatibility, performance, or a product’s failover behavior.

Resolve identity across hosts

The first group presents host-a/mpatha/W-A and host-b/mpatha/W-B. The same alias is associated with different volumes. On host-b, mpathc corresponds to W-A and is the candidate to correlate with host-a’s volume. Do not alter volume identity merely to make names match. During a rebuild, compare approved inventory with actual presentation before any writes. Equal capacity or equal LUN numbers also do not resolve a WWID mismatch. The decision should state which volume was intended and which was actually observed. If correlation remains uncertain, keep application writes disabled while storage and application owners reconcile the information.

Count surviving dependencies

In the second group, four paths combine two HBAs, two fabrics, and two controllers. Removing fabric-a leaves p3 and p4; additionally removing controller-a leaves only p4. A variant places every path on fabric-a and loses all four when that fabric fails. These results come from declared dependency sets, not from switching off equipment. The exercise requires naming the failure the design must tolerate. Adding cables within one domain may improve something else without addressing that failure. After predicting survivors, identify missing evidence: genuinely usable paths, supported workload, and recovery time. Do not present one remaining path as an automatic guarantee of meeting the SLA.

Separate login from LUN authorization

The third group receives an established session, effective IQN aps-new, and authorization containing only aps-old. The model’s access predicate remains false. This simplification discusses two separate conditions; it is not an iSCSI implementation. During diagnosis, use the mismatch to direct the next evidence collection: intended rebuilt-server identity, effective configuration, and the missing LUN’s mapping. Copying every old permission without review can introduce improper exposure. If the IQN already matches, investigation continues with the specific mapping and remaining platform conditions. Document what was confirmed and what remains unknown. A successful login is useful evidence, but it does not close analysis of the intended volume.

Interpret reservations by their specific type

The fourth group fixes host-a as holder and host-b as a registered initiator. With PR_WRITE_EXCLUSIVE, the model permits host-b reads and refuses writes. With PR_EXCLUSIVE_ACCESS, it refuses both; with PR_WRITE_EXCLUSIVE_REG_ONLY, it permits both for the registered initiator. The comparison shows why saying there is a reservation is insufficient to decide who can write. The script sends no SCSI commands, transfers no ownership, and does not test fencing. During a real intervention, confirm type, holder, registrations, and intended cluster coordination before proposing any change. Removing a reservation to bypass an error may remove necessary protection. The next step depends on the application and authorized procedure.

Decide on overlapping maintenance

Reserve ten workshop minutes for the first APS situation. One team keeps fabric-a unavailable while another wants to stop fabric-b before the window ends. Use the matrix to show that remaining paths depend on the second fabric. Produce a short decision with observed state, overlap impact, progression condition, and recovery-validation owner. If the window cannot accommodate restoring and observing fabric-a, propose rescheduling with its impact communicated. Two separate tickets do not establish technical independence. Before applying the reasoning to ONTAP, also consult the current platform, protocol, and version matrix. Failover rules vary and should not be reduced to one universal assertion.

python3 content/labs/san-evidence/run.py
# Fictional path inventory; not output from multipath -ll
# p1 hba-a fabric-a controller-a
# p2 hba-a fabric-a controller-b
# p3 hba-b fabric-b controller-a
# p4 hba-b fabric-b controller-b
# Failed {fabric-a, controller-a} -> survivor p4
# Offline model; no SAN connection or physical failure injection.
IN PRACTICE

A fictional rebuilt host displays the expected alias but a different WWID. Its iSCSI session exists and mapping still refers to the old IQN.

Common pitfalls

Confusing an alias with global identity, a session with authorization, four paths with four independent domains, or key registration with write exclusivity.

Related topics: Storage · Production Support L3 · High Availability

Take this idea with you

Before authorizing writes or a second maintenance operation, confirm the intended volume, who accesses it, and which paths survive the failures considered.

Create account

Reference: RHEL 9 DM Multipath: topology, identity, queue policy, resize and safe removal · BigSavant SAN 2026-09; selected RHEL 9, ONTAP 9 and iSCSI behavior