Hi,
I am troubleshooting a two-node Hyper-V failover cluster. During a maintenance operation, transferring CSV ownership from Node B to Node A caused a CSV operation to hang on Node A. The Cluster Service subsequently terminated and Node A was reported as isolated. Recovery involved an iDRAC power cycle.
Environment:
- Two Dell PowerEdge servers.
Windows Server 2022 Datacenter, build 20348.5622 on the affected nodes.
Dell PowerVault ME5024 shared iSCSI storage with two controllers.
Two Netgear M4350-12X12F switches in a stack.
Two dedicated 10 GbE iSCSI NICs per host, using separate subnets/VLANs, without NIC teaming.
Microsoft DSM / MPIO, Round Robin with Subset.
Eight connected iSCSI sessions per host. Previous path checks showed four Active/Optimized and four Active/Unoptimized paths per LUN.
Two CSV LUNs and a separate disk witness.
Host management, cluster communication and Live Migration share a 2 × 1 GbE LACP team.
Cluster communication is disabled on the dedicated iSCSI networks.
VM networking uses a separate 2 × 10 GbE SET switch, SwitchIndependent / HyperVPort.
Incident sequence, local time:
21:59:19 — CSV Volume2 ownership moves from Node B to Node A during node drain.
21:59:21 — Node A starts a dcm/map operation for Volume2, which does not complete.
22:06:44 — A VM resource on Volume2 exceeds its health-check timeout. The cluster restarts the corresponding RHS process.
22:07:06 — The Cluster Service on Node B stops gracefully for its restart. The problem on Node A is already present before this.
22:13:53 — Another timeout occurs for the same VM resource.
22:17:21 — After approximately 18 minutes, Node A logs the following error and its Cluster Service terminates:
FatalError: Volume 'Volume2:<redacted>' is stuck transitioning to CsvFsVolumeStateSetDownlevel (status = 121)
22:17:22 — Node B reports Node A as isolated.
Additional findings:
Node A's System log confirms Cluster Service termination with error 121.
CSV event 5120, STATUS_NO_SUCH_DEVICE, and disk reservation release errors appear around the time the Cluster Service terminates. Their timing does not establish that they initiated the incident.
No preceding iSCSI or Broadcom NIC errors were found in the reviewed System log interval.
VMMS also recorded errors referencing AVHDX files of another VM on Volume1 a few seconds before the Volume2 mapping operation became stuck. Their relevance is unclear.
No backup job was running during the incident.
After recovery, both nodes were Up, both CSVs showed Direct access on both nodes, and all eight iSCSI sessions were connected on each host.
The versions match between nodes. We have not established that either filter caused the hang.
Cluster.log records an unsuccessful attempt to capture a live dump, with status 0xd0000022. No relevant dump was found in the checked Cluster\Reports or LiveKernelReports directories.
Has anyone encountered this particular dcm/map / CsvFsVolumeStateSetDownlevel hang during CSV ownership transfer on Windows Server 2022?
Are there known issues involving CSV state transitions, these filter drivers, or this storage configuration? What tracing or dump collection would you prepare before a controlled reproduction to identify the blocked operation?Hi,
I am troubleshooting a two-node Hyper-V failover cluster. During a maintenance operation, transferring CSV ownership from Node B to Node A caused a CSV operation to hang on Node A. The Cluster Service subsequently terminated and Node A was reported as isolated. Recovery involved an iDRAC power cycle.
Environment:
Two Dell PowerEdge servers.
Windows Server 2022 Datacenter, build 20348.5622 on the affected nodes.
Dell PowerVault ME5024 shared iSCSI storage with two controllers.
Two Netgear M4350-12X12F switches in a stack.
Two dedicated 10 GbE iSCSI NICs per host, using separate subnets/VLANs, without NIC teaming.
Microsoft DSM / MPIO, Round Robin with Subset.
Eight connected iSCSI sessions per host. Previous path checks showed four Active/Optimized and four Active/Unoptimized paths per LUN.
Two CSV LUNs and a separate disk witness.
Host management, cluster communication and Live Migration share a 2 × 1 GbE LACP team.
Cluster communication is disabled on the dedicated iSCSI networks.
VM networking uses a separate 2 × 10 GbE SET switch, SwitchIndependent / HyperVPort.
Incident sequence, local time:
21:59:19 — CSV Volume2 ownership moves from Node B to Node A during node drain.
21:59:21 — Node A starts a dcm/map operation for Volume2, which does not complete.
22:06:44 — A VM resource on Volume2 exceeds its health-check timeout. The cluster restarts the corresponding RHS process.
22:07:06 — The Cluster Service on Node B stops gracefully for its restart. The problem on Node A is already present before this.
22:13:53 — Another timeout occurs for the same VM resource.
22:17:21 — After approximately 18 minutes, Node A logs the following error and its Cluster Service terminates:
FatalError: Volume 'Volume2:<redacted>' is stuck transitioning to CsvFsVolumeStateSetDownlevel (status = 121)
22:17:22 — Node B reports Node A as isolated.
Additional findings:
Node A's System log confirms Cluster Service termination with error 121.
CSV event 5120, STATUS_NO_SUCH_DEVICE, and disk reservation release errors appear around the time the Cluster Service terminates. Their timing does not establish that they initiated the incident.
No preceding iSCSI or Broadcom NIC errors were found in the reviewed System log interval.
VMMS also recorded errors referencing AVHDX files of another VM on Volume1 a few seconds before the Volume2 mapping operation became stuck. Their relevance is unclear.
No backup job was running during the incident.
After recovery, both nodes were Up, both CSVs showed Direct access on both nodes, and all eight iSCSI sessions were connected on each host.
The versions match between nodes. We have not established that either filter caused the hang.
Cluster.log records an unsuccessful attempt to capture a live dump, with status 0xd0000022. No relevant dump was found in the checked Cluster\Reports or LiveKernelReports directories.
Has anyone encountered this particular dcm/map / CsvFsVolumeStateSetDownlevel hang during CSV ownership transfer on Windows Server 2022?
Are there known issues involving CSV state transitions, these filter drivers, or this storage configuration? What tracing or dump collection would you prepare before a controlled reproduction to identify the blocked operation?