Hyper-V Failover Cluster – VMs entering Saved State after unexpected node reboot

John lamma 20 Reputation points
2026-10-08T08:50:12.9266667+00:00

Hello,

We have a 2-node Hyper-V Failover Cluster running Windows Server 2025 Datacenter and we are experiencing an issue when one of the nodes is unexpectedly restarted.

Environment

  • 2 Hyper-V nodes
  • Windows Server 2025 Datacenter
  • Failover Clustering
  • Network ATC configured for the cluster networking and Live Migration
  • Shared storage provided through an iSCSI SAN
  • VMs are hosted on the shared iSCSI storage

What works correctly

Normal Live Migration works without any problems.

We can manually migrate VMs from Node 1 to Node 2 and from Node 2 to Node 1. The migration completes successfully, without errors, and the VMs continue running normally on the destination node.

The cluster is also working normally during standard operations.

Problem

The problem occurs when we force a reboot of one of the Hyper-V nodes.

For example, if Node 1 is hosting several VMs and we force a reboot of Node 1, the cluster detects the node failure and moves the VMs to Node 2.

However, instead of all VMs starting normally on Node 2, some of them enter Saved State.

After waiting several minutes:

  • Some VMs automatically start and work normally.
  • Some VMs remain in Saved State.
  • The VMs that remain in Saved State do not automatically recover and require manual intervention.

We have checked the Event Viewer and Hyper-V/Failover Cluster logs around the time of the incident and found Event ID 1137 and Event ID 1155.

What we find strange

The important point is that Live Migration works perfectly when performed manually.

The issue only appears when a node is suddenly lost or forcefully rebooted and the cluster has to recover the VMs on the surviving node.

Because the VMs are using the same shared iSCSI storage and normal Live Migration works correctly, we are trying to understand what could be causing some VMs to enter Saved State specifically during an unexpected node failure.Hello,

We have a 2-node Hyper-V Failover Cluster running Windows Server 2025 Datacenter and we are experiencing an issue when one of the nodes is unexpectedly restarted.

Windows for business | Windows Server | Storage high availability | Clustering and high availability
0 comments No comments

Answer accepted by question author
VPHAN 45,420 Reputation points Independent Advisor
2026-10-08T09:51:21.9566667+00:00

Hi John lamma,

Your issue stems from a known timing race condition involving how Windows Server 2025 handles ungraceful failovers with iSCSI storage. When a node crashes unexpectedly, the surviving node attempts to bring the virtual machines online but must wait for the iSCSI SAN to release the stale storage locks held by the dead node. If this lock transfer exceeds the internal cluster timeout threshold, the start sequence fails. To protect your virtual disks from data corruption, Hyper-V defensively drops the affected virtual machines into a Saved State. This timeout generates Event ID 1137, which indicates a cluster resource failed to come online, and Event ID 1155, indicating the pending move operation could not complete. Manual live migrations succeed perfectly because storage locks are transferred gracefully between nodes without triggering these timeouts.

Windows Server 2025 is a newly released operating system, so you must ensure both nodes have the absolute latest Cumulative Updates installed. Microsoft continually integrates cluster recovery and timing fixes into these monthly rollups. If your servers are fully patched and the issue persists, you are encountering a known bug with ungraceful failovers and will need to wait for a subsequent official update to permanently resolve this behavior.

In the meantime, you can confirm this exact storage lock timeout by generating a cluster log via PowerShell. This command compiles the raw diagnostic traces into a readable text document located in the C:\Windows\Cluster\Reports directory on your local system drive. By reviewing the exact timestamps of a forced node failure within that log, you can see the precise millisecond the virtual machine management service gave up waiting for the iSCSI lock and defaulted to the Saved State.

Hope this answer has brought you some useful information. If it did, please hit “accept answer”. Should you have any questions, feel free to leave a comment.

VPHAN

Was this answer helpful?

1 person found this answer helpful.
0 comments No comments

1 additional answer

Sort by: Most helpful
  1. AI answer

    2026-10-08T08:50:52.72+00:00

    The pattern points to failover recovery behavior, not a general Live Migration problem. Manual Live Migration succeeding means the shared iSCSI storage path and normal migration path are working. When the node is lost unexpectedly, the cluster must recover VMs after an unclean interruption, and that can leave some VMs in Saved or Paused states instead of bringing them fully online.

    A practical troubleshooting path is:

    1. Validate cluster health first Run Test-Cluster and review the results. Cluster configuration issues can cause failed live migrations, unexpected VM restarts, and unexpected failover events.
    2. Review the exact VM state after failover In Hyper-V Manager, Failover Cluster Manager, or PowerShell, identify which VMs remain in Saved, Paused, Restoring, or other transitional states. The troubleshooting guidance specifically calls out these states as important indicators.
    3. Check for cluster metadata or configuration issues Improper failovers and unsynchronized cluster states can cause cluster role or metadata corruption. Recommended fixes are:
      • Use Failover Cluster Manager for VM operations.
      • Perform live migrations to force metadata synchronization and clear warnings.
      • If needed, verify VM and storage path settings, including disk and network adapter assignments.
    4. Check VM configuration version and VM configuration integrity After migrations or upgrades, a VM configuration version mismatch can cause issues. In Hyper-V Manager, right-click the VM and select Upgrade Configuration Version, or run:
         Update-VMVersion
      
      Also verify that the VM folder still contains the expected files: .vmcx, .vmrs, .vhdx, and .avhdx.
    5. If a VM is stuck in Saved state, clean up saved-state files carefully The guidance for backup/checkpoint/storage failures says to delete .vmrs and .vmgs files in the VM folder if the VM is stuck in a Saved state. This is a potentially destructive action because saved state data is discarded. Use it only for affected VMs after confirming the VM disks are intact.
    6. Check storage accessibility on each node Since the storage is iSCSI-backed, verify the shared volume path from each node and confirm the VM files are accessible. If a disk volume is RAW or missing, reassign its drive letter and run:
         chkdsk <DriveLetter>: /f /r
      
      Also review whether any file or path is locked by another process; restarting the node can release stale locks.
    7. Review cluster resource state and failover side effects Inspect cluster resources with PowerShell such as Get-ClusterResource. The guidance also notes that improper maintenance, such as rebooting cluster nodes without draining roles or placing them in maintenance mode, is an operational cause of cluster and VM issues.
    8. Collect logs if the issue repeats Gather:
      • Cluster logs: Get-ClusterLog -UseLocalTime -Destination <Path>
      • Hyper-V VMMS/Admin log
      • Hyper-V Worker/Admin log
      • VM configuration details such as Get-VM * | select Name, Version

    What this most strongly suggests from the documented causes:

    • Improper failover / unsynchronized cluster metadata after an abrupt node loss
    • Saved-state cleanup or stale runtime state on affected VMs
    • Storage path or lock issues that appear only during crash recovery, not during clean manual Live Migration
    • VM configuration version or configuration file issues on the VMs that consistently remain in Saved State

    One important operational note: the guidance explicitly says to drain roles before rebooting a node:

    Suspend-ClusterNode -Name <NodeName> -Drain
    

    That does not fix an unexpected crash, but it does explain why clean manual movement works while forced reboot recovery does not.

    If the affected VMs instead show as Paused together with Event IDs 1135, 1795, or 5257, the documented resolution is to restart those VMs from Hyper-V Manager or Failover Cluster Manager and monitor them until they return to an operational state.


    References

    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.