Hi John lamma,
Your issue stems from a known timing race condition involving how Windows Server 2025 handles ungraceful failovers with iSCSI storage. When a node crashes unexpectedly, the surviving node attempts to bring the virtual machines online but must wait for the iSCSI SAN to release the stale storage locks held by the dead node. If this lock transfer exceeds the internal cluster timeout threshold, the start sequence fails. To protect your virtual disks from data corruption, Hyper-V defensively drops the affected virtual machines into a Saved State. This timeout generates Event ID 1137, which indicates a cluster resource failed to come online, and Event ID 1155, indicating the pending move operation could not complete. Manual live migrations succeed perfectly because storage locks are transferred gracefully between nodes without triggering these timeouts.
Windows Server 2025 is a newly released operating system, so you must ensure both nodes have the absolute latest Cumulative Updates installed. Microsoft continually integrates cluster recovery and timing fixes into these monthly rollups. If your servers are fully patched and the issue persists, you are encountering a known bug with ungraceful failovers and will need to wait for a subsequent official update to permanently resolve this behavior.
In the meantime, you can confirm this exact storage lock timeout by generating a cluster log via PowerShell. This command compiles the raw diagnostic traces into a readable text document located in the C:\Windows\Cluster\Reports directory on your local system drive. By reviewing the exact timestamps of a forced node failure within that log, you can see the precise millisecond the virtual machine management service gave up waiting for the iSCSI lock and defaulted to the Saved State.
Hope this answer has brought you some useful information. If it did, please hit “accept answer”. Should you have any questions, feel free to leave a comment.
VPHAN