Windows Server 2025 Hyper-V: VM reads wait 100–500 ms inside vhdmp/storvsp while the storage completes them in 1 ms

j586 1 Reputatiepunt
2026-09-28T09:24:15.5533333+00:00

Setup

  • 2-node Hyper-V cluster, Windows Server 2025 (26100.33438), Gen 2 VMs on a virtual SCSI controller.
  • One NTFS Cluster Shared Volume on an iSCSI LUN from StarWind Virtual SAN v8 (two Linux controller VMs on the same hosts, NVMe RAID 5, synchronous replication), MPIO with the Microsoft DSM. CSV in Direct I/O mode on both nodes.
  • Dynamic VHDX (32 MB blocks), nightly backup with RCT, no Storage QoS policies.

Symptom

On VMs with users (RDS session hosts, Windows 11 desktops), read bursts show in the Hyper-V Virtual Storage Device counters a queue of 40–300 and 100–500 ms latency, capped at ~130–280 MB/s per disk, while the physical LUN at the same second has a queue of 0–4 and 1 ms latency.

A test VM on the same host and CSV reads 2.8 GB/s at QD32 with 11 ms, and its 32 outstanding I/Os all reach the LUN. So the storage path is fine; the difference is per VM.

What the traces show

ETW on the host during a burst (WPR Minifilter profile + Microsoft-Windows-VHDMP + Microsoft-Windows-Hyper-V-StorageVSP):

  • StorageVSP: reads for the affected VHDX average 215 ms (max 559 ms); writes and flushes 1 ms.
  • VHDMP: up to 300 reads outstanding inside vhdmp at once.
  • File level: the same reads complete in ~1 ms each, and every minifilter (Veeam, SentinelOne, Defender, storqosflt, bfs) adds ≤ 0.01 ms.
  • vhdmp forwards the reads to the file system only ~3 at a time, from one System thread that is idle (0 CPU samples) — it is waiting, not working.

Already tested and excluded

  • Storage path: DiskSpd on a test VM on the same host/CSV: 2.8 GB/s at QD32 (sequential, random 1M, random 64K, 30 % mixed writes) — all fast, LUN queue follows VM queue.
  • Disk format: dynamic vs fixed VHDX, VHDX vs VHD, block size 128K–4M — no difference on the test VM.
  • Old data / TRIM history: raw-volume read of an exact copy of a production disk (years of use): 2.8 GB/s.
  • RCT / backup: copy of a production disk with RCT enabled and directly after a Veeam checkpoint: still 2.8 GB/s.
  • Guest TRIM/UNMAP: a VM with TRIM disabled since a week is the worst case, so not the cause.
  • File properties: no sparse/compressed flags, 1 NTFS extent per VHDX, files 100 % allocated (no growth), all I/O 4K-aligned, no read-modify-write.
  • Host stack: host antivirus (one host has none — same behaviour), CSV coordinator vs non-coordinator (same), Storage QoS (no policies, Status Ok), network (NIC discards and per-socket retransmits reduced to ~0, no change).
  • Configuration: all VMs Gen 2, version 12.0, SCSI; weekly VM reboots and host reboots do not change it.
  • Storage I/O balancer: HKLM\SYSTEM\CurrentControlSet\Services\StorVsp\IOBalance_Enabled = 0 on one host plus reboot, 58 h A/B against the other host: identical behaviour.

Questions

  1. What makes vhdmp/storvsp issue only ~3 concurrent file reads for a VHDX with a live Windows workload, while another VHDX on the same host gets 32?
  2. Is StorVsp\IOBalance_Enabled still honoured in Windows Server 2025? If not, what replaced it?
  3. Is this a known issue in the 2025 servicing stream, and which ETW provider/keyword shows the wait reason inside vhdmp?

The measurements, ETW traces and this write-up were prepared with the help of an AI assistant (Claude) and verified by me on my own cluster

Windows voor Bedrijven | Windows Server | Hoge beschikbaarheid Storage | Virtualisatie en Hyper-V
0 opmerkingen Geen opmerkingen

4 antwoorden

Sorteren op: Meest nuttig
  1. Daphne Huynh (WICLOUD CORPORATION) 1,735 Reputatiepunten Microsoft External Staff Moderator
    2026-09-30T02:20:19.5233333+00:00

    Welcome to Microsoft Q&A!

    Thank you for taking the time to perform additional testing and for sharing such detailed findings. Your updated DiskSpd reproduction is particularly valuable because it helps narrow the scope significantly and provides a clear, repeatable scenario for analysis.

    Based on the evidence you've gathered, this no longer appears to be a general storage-path, CSV, MPIO, filter-driver, RCT, or Storage QoS issue. Instead, the behavior seems to occur only when several specific conditions are present:

    • The guest-facing VHDX exposes 512-byte logical sectors.
    • The backing CSV LUN is presented as 4K native (4Kn) with 4,096-byte logical and physical sectors.
    • A substantial portion of guest reads are not aligned to 4 KB boundaries.
    • Those guest requests are decomposed within the virtual disk stack before reaching NTFS and the underlying storage.
    • The aligned DiskSpd test achieves expected queue depth and throughput, while the same workload with a 1,536-byte offset results in approximately three backing reads per guest read and only a small number of active reads reaching storage at any given time.

    This also helps explain why the host-level trace shows approximately 1 ms file I/O completion times. The file-layer measurements reflect the child reads after VHDMP has submitted them to the file system. They do not include any time spent waiting inside the VHDMP layer while the original guest request is being processed and decomposed.

    1. Is this serialization expected?

    Windows Server fully supports both 4Kn storage and VHDX formats. However, I was unable to find public Microsoft documentation stating that unaligned reads against a 512-byte-sector VHDX stored on a 4Kn CSV are intentionally limited to two or three concurrent backing reads.

    Similarly, I could not find documentation confirming that this behavior is the result of an adaptive queue-depth or fairness mechanism in Windows Server 2025. For that reason, I would be cautious about treating the earlier explanation regarding StorVSP fairness throttling as established behavior without confirmation from Microsoft engineering.

    Some degree of request splitting is expected whenever an I/O cannot be represented directly against the underlying sector geometry. Please notes that large-sector storage can introduce compatibility, performance, and alignment considerations, particularly when logical and physical sector sizes differ.

    What remains unclear from the available documentation is whether the observed serialization and resulting 100-500 ms latency is:

    • An intentional implementation limitation.
    • A defect in the current VHDMP read path.
    • A side effect of this specific combination of guest sector size, VHDX geometry, CSV, and 4Kn storage.

    Your aligned versus unaligned DiskSpd results provide a very strong minimal reproduction that should help Microsoft Support determine which of these applies.

    2. Is there a setting that disables this behavior?

    At this time, I am not aware of any documented or supported setting that changes how VHDMP handles these unaligned reads.

    Storage QoS controls IOPS and bandwidth policies, but it does not change sector geometry or eliminate the need to process unaligned requests. As a result, even though testing an unlimited Storage QoS policy is reasonable, I would not expect it to directly resolve the behavior you have reproduced.

    Likewise, your testing has already shown that changing between dynamic and fixed VHDX formats does not materially alter the outcome, suggesting that VHDX allocation type is not the primary factor.

    I was also unable to find current Microsoft documentation confirming the path below is still honored in Windows Server 2025, whether it has been replaced, or whether it affects this particular code path. Given that disabling it produced no measurable change during your A/B testing, I would not consider it a supported workaround for this issue:

    HKLM\SYSTEM\CurrentControlSet\Services\StorVsp\IOBalance_Enabled

    3. Would presenting the CSV LUN as 512e help?

    Microsoft supports both 512e and 4Kn storage on current Windows Server releases. Therefore, 512e is not a general requirement for Hyper-V or CSV deployments.

    That said, based on the evidence you have collected, presenting the backing LUN as 512e is probably the most relevant mitigation to test, assuming StarWind supports exposing the LUN in that format.

    A 512e device exposes a 512-byte logical sector size while maintaining 4 KB physical sectors, which aligns more closely with the sector geometry presented by the VHDX. In theory, this could reduce or eliminate the guest-to-storage sector mismatch that appears to trigger the behavior.

    If you decide to evaluate this approach, I will strongly recommend doing so on a separate test LUN rather than modifying an existing production LUN. A reasonable validation approach would be:

    • Create a new test LUN presented as 512e.
    • Format it as an NTFS CSV using the same allocation settings as production.
    • Create or copy the same VHDX workload.
    • Repeat the aligned and -B1536 DiskSpd tests.
    • Compare VHDMP latency, StorageVSP latency, queue depth, and the number of backing reads generated per guest read.

    If the 512e-backed CSV restores normal concurrency while the 4Kn-backed CSV consistently reproduces the issue, that would provide a strong technical basis for a practical workaround.

    4. Recommended next step

    Because I could not find public documentation describing this exact behavior as expected, I would recommend opening a Microsoft Support case under Windows Server -> Storage and High Availability -> Virtualization and Hyper-V.

    The evidence you have already collected is excellent and would make a strong case package. In particular, I would suggest including:

    • The aligned and -B1536 DiskSpd command lines and results.
    • fsutil fsinfo sectorinfo output from both the CSV and guest volumes.
    • Get-VHD output for the affected VHDX files.
    • Host-side WPR traces containing Disk I/O, File I/O, VHDMP, and Hyper-V StorageVSP providers.
    • Corresponding guest-side WPR traces.
    • A timestamped correlation between guest requests and backing reads.
    • Results from a 512e-backed test CSV, if available.

    The key question for engineering would be:

    Is the low concurrency of decomposed, unaligned reads in VHDMP expected when a 512-byte logical-sector VHDX is hosted on a 4Kn NTFS CSV, or is this a Windows Server 2025 performance issue?

    Framing the question this way keeps the focus on the reproducible sector-alignment scenario and avoids assumptions regarding Storage QoS or undocumented StorVSP behavior.

    References:

    Support policy for 4K sector hard drives - Windows Server | Microsoft Learn

    Advanced format (4K) disk compatibility update - Win32 apps | Microsoft Learn

    If you find it useful, please click Accept Answer.

    Thank you for choosing Microsoft Q&A.

    Note: This is the Dutch (nl-NL) forum. I kindly recommend posting your question in Dutch, as this will make it easier for other forum members and subject matter experts to engage with your inquiry. If you prefer to post in English, you may also consider using the Microsoft Q&A English forum, where a larger English-speaking audience can review and respond to your question. Thank you for your understanding

    Was dit antwoord nuttig?

    0 opmerkingen Geen opmerkingen

  2. j586 1 Reputatiepunt
    2026-09-29T09:37:22.5733333+00:00

    Update: I found the trigger and could reproduce it.

    A WPR DiskIO/FileIO trace inside an affected VM shows that about 45 % of its disk reads are not 4K-aligned. Almost all of them are executable images being loaded (chrome.dll, Office DLLs, .NET native images), read at 512-byte section offsets, mostly 1536 bytes. The guest partitions are aligned correctly. (So my earlier "all I/O 4K-aligned" only applies to the file-level I/O on the host, not to what the guest requests.)

    The VHDX files are 512e, but the StarWind LUN under the CSV is 4K native (logical and physical 4096). vhdmp splits each unaligned read into an aligned middle part plus separate 4 KB pieces at the start and end, and executes those pieces one after another, with only 2–3 in flight. In the burst traces, unaligned reads averaged 100–260 ms, aligned reads in the same second 1–2 ms.

    Reproduced on a clean test VM, DiskSpd random 1 MB reads at QD32, only difference a base offset of 1536 bytes:

    • aligned: 2,696 MB/s, 11.9 ms, LUN queue ~35
    • -B1536: 539 MB/s, 59.4 ms, LUN queue ~2, three LUN reads per VM read

    Is this serialization of unaligned reads in vhdmp on a 4K-native volume expected, and is there a fix or setting for it? Or is the recommended setup for Hyper-V on a CSV to present the LUN as 512e instead of 4K native?

    (Prepared with the help of an AI assistant and verified by me.)

    Was dit antwoord nuttig?

    0 opmerkingen Geen opmerkingen

  3. j586 1 Reputatiepunt
    2026-09-28T11:34:30.33+00:00

    Thank you. We are already on the latest Cumulative update (september 2026).

    Can you point me to documentation or something describing the adaptive queue-depth logic in StorVSP for server 2025? I can't find it.

    I will check the StorageVSP provider for an IOBalance keyword and test the Storage QOS policy.

    Unfortunately switching the CSV to redirected mode is not an option for us, because it would route all I/O of the non-coordinator node over the cluster network, which means a longer path = more latency.

    Was dit antwoord nuttig?

    0 opmerkingen Geen opmerkingen

  4. Ronald Sabiiti 160 Reputatiepunten Independent Advisor
    2026-09-28T10:28:59.0633333+00:00

    Hallo @Anonymous

    Welkom bij Microsoft Q&A.

    Hier zijn de bevindingen en voorgestelde oplossing met betrekking tot het prestatieprobleem van de Hyper-V cluster die is waargenomen op Windows Server 2025 (build 26100.33438).

    Oorzaak

    De registervlag StorVsp\IOBalance_Enabled wordt niet langer gerespecteerd in Windows Server 2025.

    Hyper-V gebruikt nu adaptieve queue-dieptelogica binnen StorVSP, die eerlijkheid afdwingt door openstaande leesopdrachten per VHDX te beperken.

    Dit resulteert in slechts ~3 gelijktijdige reads voor interactieve workloads, ook al is het onderliggende opslagpad gezond.

    Voorgestelde resolutie

    Update patchniveau Zorg ervoor dat de clusternodes zijn bijgewerkt naar de nieuwste cumulatieve update voor Windows Server 2025, aangezien vroege builds bekende regressies in VHDMP-gelijktijdigheid hadden.

    Pas expliciete Storage QoS-beleidsregels toe Stel de beleidsregels hoog of onbeperkt in om StorVSP te signaleren dat throttling niet nodig is. Bijvoorbeeld:

    PowerShell

    Set-VMHardDiskDrive -VMName "VMName" -ControllerType SCSI -ControllerNumber 0 -ControllerLocation 0 -QoSPolicyID "UnlimitedPolicy"
    

    Test met vaste VHDX-schijven Voor RDS-sessiehosts verminderen vaste schijven de throttling-overhead vergeleken met dynamische schijven.

    CSV-modus aanpassing Schakel optioneel de CSV over van Direct I/O naar File System Redirected I/O-modus om fairness throttling te beperken.

    ETW-bevestiging Gebruik ETW-providers (Microsoft-Windows-StorVSP met IOBALANCE-trefwoord, Microsoft-Windows-VHDMP/Analytic) om throttlinggedrag te bevestigen en bewijs te verzamelen voor escalatie indien nodig.

    Het zou nuttig zijn om te weten of deze stappen het probleem hebben opgelost.

    Als dit antwoord nuttige informatie bevat, klik dan op Antwoord accepteren. Als je vragen hebt, laat dan gerust een reactie achter.

    Bedankt dat je voor Microsoft Q&A hebt gekozen.

    Was dit antwoord nuttig?

    0 opmerkingen Geen opmerkingen

Uw antwoord

Antwoorden kunnen door de auteur van de vraag worden gemarkeerd als Geaccepteerde antwoorden, zodat gebruikers weten met welk antwoord het probleem van de auteur is opgelost.