Windows Server 2022 Hyper-V cluster: CSV ownership transfer hangs in CsvFsVolumeStateSetDownlevel, then node becomes isolated

Marcin Rabiniak 0 Reputation points
2026-09-25T07:51:27.7366667+00:00

Hi,

I am troubleshooting a two-node Hyper-V failover cluster. During a maintenance operation, transferring CSV ownership from Node B to Node A caused a CSV operation to hang on Node A. The Cluster Service subsequently terminated and Node A was reported as isolated. Recovery involved an iDRAC power cycle.

Environment:

  • Two Dell PowerEdge servers.

Windows Server 2022 Datacenter, build 20348.5622 on the affected nodes.

Dell PowerVault ME5024 shared iSCSI storage with two controllers.

Two Netgear M4350-12X12F switches in a stack.

Two dedicated 10 GbE iSCSI NICs per host, using separate subnets/VLANs, without NIC teaming.

Microsoft DSM / MPIO, Round Robin with Subset.

Eight connected iSCSI sessions per host. Previous path checks showed four Active/Optimized and four Active/Unoptimized paths per LUN.

Two CSV LUNs and a separate disk witness.

Host management, cluster communication and Live Migration share a 2 × 1 GbE LACP team.

Cluster communication is disabled on the dedicated iSCSI networks.

VM networking uses a separate 2 × 10 GbE SET switch, SwitchIndependent / HyperVPort.

Incident sequence, local time:

21:59:19 — CSV Volume2 ownership moves from Node B to Node A during node drain.

21:59:21 — Node A starts a dcm/map operation for Volume2, which does not complete.

22:06:44 — A VM resource on Volume2 exceeds its health-check timeout. The cluster restarts the corresponding RHS process.

22:07:06 — The Cluster Service on Node B stops gracefully for its restart. The problem on Node A is already present before this.

22:13:53 — Another timeout occurs for the same VM resource.

22:17:21 — After approximately 18 minutes, Node A logs the following error and its Cluster Service terminates:

FatalError: Volume 'Volume2:<redacted>' is stuck transitioning to CsvFsVolumeStateSetDownlevel (status = 121)

22:17:22 — Node B reports Node A as isolated.

Additional findings:

Node A's System log confirms Cluster Service termination with error 121.

CSV event 5120, STATUS_NO_SUCH_DEVICE, and disk reservation release errors appear around the time the Cluster Service terminates. Their timing does not establish that they initiated the incident.

No preceding iSCSI or Broadcom NIC errors were found in the reviewed System log interval.

VMMS also recorded errors referencing AVHDX files of another VM on Volume1 a few seconds before the Volume2 mapping operation became stuck. Their relevance is unclear.

No backup job was running during the incident.

After recovery, both nodes were Up, both CSVs showed Direct access on both nodes, and all eight iSCSI sessions were connected on each host.

The versions match between nodes. We have not established that either filter caused the hang.

Cluster.log records an unsuccessful attempt to capture a live dump, with status 0xd0000022. No relevant dump was found in the checked Cluster\Reports or LiveKernelReports directories.

Has anyone encountered this particular dcm/map / CsvFsVolumeStateSetDownlevel hang during CSV ownership transfer on Windows Server 2022?

Are there known issues involving CSV state transitions, these filter drivers, or this storage configuration? What tracing or dump collection would you prepare before a controlled reproduction to identify the blocked operation?Hi,

I am troubleshooting a two-node Hyper-V failover cluster. During a maintenance operation, transferring CSV ownership from Node B to Node A caused a CSV operation to hang on Node A. The Cluster Service subsequently terminated and Node A was reported as isolated. Recovery involved an iDRAC power cycle.

Environment:

Two Dell PowerEdge servers.

Windows Server 2022 Datacenter, build 20348.5622 on the affected nodes.

Dell PowerVault ME5024 shared iSCSI storage with two controllers.

Two Netgear M4350-12X12F switches in a stack.

Two dedicated 10 GbE iSCSI NICs per host, using separate subnets/VLANs, without NIC teaming.

Microsoft DSM / MPIO, Round Robin with Subset.

Eight connected iSCSI sessions per host. Previous path checks showed four Active/Optimized and four Active/Unoptimized paths per LUN.

Two CSV LUNs and a separate disk witness.

Host management, cluster communication and Live Migration share a 2 × 1 GbE LACP team.

Cluster communication is disabled on the dedicated iSCSI networks.

VM networking uses a separate 2 × 10 GbE SET switch, SwitchIndependent / HyperVPort.

Incident sequence, local time:

21:59:19 — CSV Volume2 ownership moves from Node B to Node A during node drain.

21:59:21 — Node A starts a dcm/map operation for Volume2, which does not complete.

22:06:44 — A VM resource on Volume2 exceeds its health-check timeout. The cluster restarts the corresponding RHS process.

22:07:06 — The Cluster Service on Node B stops gracefully for its restart. The problem on Node A is already present before this.

22:13:53 — Another timeout occurs for the same VM resource.

22:17:21 — After approximately 18 minutes, Node A logs the following error and its Cluster Service terminates:

FatalError: Volume 'Volume2:<redacted>' is stuck transitioning to CsvFsVolumeStateSetDownlevel (status = 121)

22:17:22 — Node B reports Node A as isolated.

Additional findings:

Node A's System log confirms Cluster Service termination with error 121.

CSV event 5120, STATUS_NO_SUCH_DEVICE, and disk reservation release errors appear around the time the Cluster Service terminates. Their timing does not establish that they initiated the incident.

No preceding iSCSI or Broadcom NIC errors were found in the reviewed System log interval.

VMMS also recorded errors referencing AVHDX files of another VM on Volume1 a few seconds before the Volume2 mapping operation became stuck. Their relevance is unclear.

No backup job was running during the incident.

After recovery, both nodes were Up, both CSVs showed Direct access on both nodes, and all eight iSCSI sessions were connected on each host.

The versions match between nodes. We have not established that either filter caused the hang.

Cluster.log records an unsuccessful attempt to capture a live dump, with status 0xd0000022. No relevant dump was found in the checked Cluster\Reports or LiveKernelReports directories.

Has anyone encountered this particular dcm/map / CsvFsVolumeStateSetDownlevel hang during CSV ownership transfer on Windows Server 2022?

Are there known issues involving CSV state transitions, these filter drivers, or this storage configuration? What tracing or dump collection would you prepare before a controlled reproduction to identify the blocked operation?

Windows for business | Windows Server | Storage high availability | Clustering and high availability
0 comments No comments

2 answers

Sort by: Most helpful
  1. Allan Solomon Mejia 10,225 Reputation points
    2026-09-25T13:56:08.8033333+00:00

    Hello @Marcin Rabiniak

    Your timeline points to the CSV ownership/state transition on Node A as the point to investigate:

    CSV ownership moves to Node A → dcm/map starts → operation hangs → CsvFsVolumeStateSetDownlevel times out

    Each CSV has a coordinator node that owns the underlying physical disk resource, even though all cluster nodes can access the CSV concurrently. CSV coordinator ownership can also be moved between nodes.

    Because the failure consistently occurs during that transition, focus controlled reproduction on what blocks the CSV state change rather than assuming Event 5120 or the later node isolation caused the problem.

    Before reproducing, capture:

    Get-ClusterSharedVolume
    Get-ClusterLog -UseLocal -TimeSpan 30 -Destination C:\ClusterLogs
    

    Get-ClusterSharedVolume confirms the CSV coordinator/owner, while cluster logs are a primary source for diagnosing Failover Clustering problems. (learn.microsoft.com)

    For reproduction, this is the supported cmdlet for explicitly changing CSV coordinator ownership:

    Move-ClusterSharedVolume -Name "<CSV name>" -Node "<NodeA>"
    

    Correlate the resulting Cluster.log with the FailoverClustering/Operational, System, storage/iSCSI/MPIO, and CSV events around the exact ownership-transfer timestamp.

    The later STATUS_NO_SUCH_DEVICE, reservation-release errors, Event 5120, and node isolation are important, but based on your timeline, they occur after the dcm/map operation is already stuck, so they don't establish the initiating cause.

    I cannot find verified Microsoft documentation identifying CsvFsVolumeStateSetDownlevel with status 121 as a known Windows Server 2022 defect, or documentation tying this exact failure to ME5024, Microsoft DSM/MPIO configuration, or the filter drivers you mentioned.

    Given that the cluster attempted but failed to capture a live dump, if the issue is reproducible, involve Microsoft Support before another production reproduction so they can configure the appropriate hang/dump tracing in advance. The blocked operation needs to be captured while it occurs to determine whether the wait is in CSVFS, storage/MPIO, a filter driver, or another component.

    References:

    Cluster Shared Volumes overview

    Move-ClusterSharedVolume


    Help make this community better for everyone: If this answer helped or resolved your issue, please accept it or upvote it. If not, share more details in a comment so we can continue the discussion and find the right solution. Thank you.

    Was this answer helpful?

    0 comments No comments

  2. AI answer

    2026-09-25T07:52:02.38+00:00

    The symptoms align with a storage/CSV access failure rather than a confirmed cluster-network isolation root cause.

    Supported indicators from the described timeline are:

    1. CSV access failure pattern
      • Event IDs 5120/5142 are associated with CSV paused or disconnected states and VM disk I/O errors.
      • STATUS_NO_SUCH_DEVICE and reservation-release failures around the same time are consistent with storage path or disk visibility problems.
    2. Likely areas to validate first
      • Storage pathing / MPIO: confirm all paths are online by using mpclaim -s -d.
      • Disk health / visibility: review Get-PhysicalDisk | Format-Table and check for storage-related events such as 157, 11, 129, and 153.
      • Drivers / firmware / DSM-MPIO stack: update storage drivers, SAN firmware, and DSM/MPIO software.
      • CSV/network transport factors: verify there is no network bottleneck or misconfiguration affecting CSV traffic, and confirm adapter configuration is aligned across all nodes.
      • Filter drivers / antivirus: inspect filter drivers with fltmc; remove unsupported filter drivers and avoid multiple antivirus products.
    3. How to interpret Event ID 5120
      • If Event ID 5120 shows STATUS_CLUSTER_CSV_AUTO_PAUSE_ERROR or c0130021, that specific event can be ignored.
      • If Event ID 5120 shows other error codes, it indicates a real problem and the error code in each event should be reviewed.
    4. Data to collect before a controlled reproduction Collect the following before reproducing the ownership transfer:
      • Cluster logs:
             Get-ClusterLog -Destination \\<FolderPath> -UseLocal -TimeSpan <Minutes>
        
      • Export System, Application, and FailoverClustering event logs.
      • Network trace:
             Netsh trace start scenario=GENERAL capture=yes tracefile=<path>
        
      • Process dump files as needed by using Sysinternals or built-in Windows tools.
      • Storage and hardware diagnostics from the storage vendor tools.
      • Security software logs if any security/filtering product is installed.
    5. If the cluster service or disk resource is hanging For failover cluster hangs, triage the Windows Error Reporting data under: C:\ProgramData\Microsoft\Windows\WER\ReportArchive\... Start with Report.wer. Relevant dumps can include RHS, clussvc.exe, and the process hosting smphost. For physical disk hangs, the DumpServicesSmphost property controls which services/processes are collected.
    6. Recovery actions if a node remains quarantined after reproduction General recovery guidance includes:
      • Start-ClusterNode -ClearQuarantine
      • restart and monitor logs
      • if necessary, evict and re-add the node
    7. Important caution Commands that remove cluster resources or alter storage pool ownership are destructive in the wrong scenario. Avoid using CSV removal or storage-pool recovery steps unless the environment actually matches that recovery article and there is a validated rollback plan.

    The available material does not document a specific known Windows Server 2022 bug for CsvFsVolumeStateSetDownlevel, a specific issue with the named filter drivers, or a Dell ME5024-specific defect. The supported conclusion is to treat this as a CSV/storage-path investigation and collect cluster, event, network, dump, vendor storage, and filter-driver evidence around the ownership transfer.


    References:

    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.