Welcome to Microsoft Q&A!
Thank you for taking the time to perform additional testing and for sharing such detailed findings. Your updated DiskSpd reproduction is particularly valuable because it helps narrow the scope significantly and provides a clear, repeatable scenario for analysis.
Based on the evidence you've gathered, this no longer appears to be a general storage-path, CSV, MPIO, filter-driver, RCT, or Storage QoS issue. Instead, the behavior seems to occur only when several specific conditions are present:
- The guest-facing VHDX exposes 512-byte logical sectors.
- The backing CSV LUN is presented as 4K native (4Kn) with 4,096-byte logical and physical sectors.
- A substantial portion of guest reads are not aligned to 4 KB boundaries.
- Those guest requests are decomposed within the virtual disk stack before reaching NTFS and the underlying storage.
- The aligned DiskSpd test achieves expected queue depth and throughput, while the same workload with a 1,536-byte offset results in approximately three backing reads per guest read and only a small number of active reads reaching storage at any given time.
This also helps explain why the host-level trace shows approximately 1 ms file I/O completion times. The file-layer measurements reflect the child reads after VHDMP has submitted them to the file system. They do not include any time spent waiting inside the VHDMP layer while the original guest request is being processed and decomposed.
1. Is this serialization expected?
Windows Server fully supports both 4Kn storage and VHDX formats. However, I was unable to find public Microsoft documentation stating that unaligned reads against a 512-byte-sector VHDX stored on a 4Kn CSV are intentionally limited to two or three concurrent backing reads.
Similarly, I could not find documentation confirming that this behavior is the result of an adaptive queue-depth or fairness mechanism in Windows Server 2025. For that reason, I would be cautious about treating the earlier explanation regarding StorVSP fairness throttling as established behavior without confirmation from Microsoft engineering.
Some degree of request splitting is expected whenever an I/O cannot be represented directly against the underlying sector geometry. Please notes that large-sector storage can introduce compatibility, performance, and alignment considerations, particularly when logical and physical sector sizes differ.
What remains unclear from the available documentation is whether the observed serialization and resulting 100-500 ms latency is:
- An intentional implementation limitation.
- A defect in the current VHDMP read path.
- A side effect of this specific combination of guest sector size, VHDX geometry, CSV, and 4Kn storage.
Your aligned versus unaligned DiskSpd results provide a very strong minimal reproduction that should help Microsoft Support determine which of these applies.
2. Is there a setting that disables this behavior?
At this time, I am not aware of any documented or supported setting that changes how VHDMP handles these unaligned reads.
Storage QoS controls IOPS and bandwidth policies, but it does not change sector geometry or eliminate the need to process unaligned requests. As a result, even though testing an unlimited Storage QoS policy is reasonable, I would not expect it to directly resolve the behavior you have reproduced.
Likewise, your testing has already shown that changing between dynamic and fixed VHDX formats does not materially alter the outcome, suggesting that VHDX allocation type is not the primary factor.
I was also unable to find current Microsoft documentation confirming the path below is still honored in Windows Server 2025, whether it has been replaced, or whether it affects this particular code path. Given that disabling it produced no measurable change during your A/B testing, I would not consider it a supported workaround for this issue:
HKLM\SYSTEM\CurrentControlSet\Services\StorVsp\IOBalance_Enabled
3. Would presenting the CSV LUN as 512e help?
Microsoft supports both 512e and 4Kn storage on current Windows Server releases. Therefore, 512e is not a general requirement for Hyper-V or CSV deployments.
That said, based on the evidence you have collected, presenting the backing LUN as 512e is probably the most relevant mitigation to test, assuming StarWind supports exposing the LUN in that format.
A 512e device exposes a 512-byte logical sector size while maintaining 4 KB physical sectors, which aligns more closely with the sector geometry presented by the VHDX. In theory, this could reduce or eliminate the guest-to-storage sector mismatch that appears to trigger the behavior.
If you decide to evaluate this approach, I will strongly recommend doing so on a separate test LUN rather than modifying an existing production LUN. A reasonable validation approach would be:
- Create a new test LUN presented as 512e.
- Format it as an NTFS CSV using the same allocation settings as production.
- Create or copy the same VHDX workload.
- Repeat the aligned and -B1536 DiskSpd tests.
- Compare VHDMP latency, StorageVSP latency, queue depth, and the number of backing reads generated per guest read.
If the 512e-backed CSV restores normal concurrency while the 4Kn-backed CSV consistently reproduces the issue, that would provide a strong technical basis for a practical workaround.
4. Recommended next step
Because I could not find public documentation describing this exact behavior as expected, I would recommend opening a Microsoft Support case under Windows Server -> Storage and High Availability -> Virtualization and Hyper-V.
The evidence you have already collected is excellent and would make a strong case package. In particular, I would suggest including:
- The aligned and -B1536 DiskSpd command lines and results.
- fsutil fsinfo sectorinfo output from both the CSV and guest volumes.
- Get-VHD output for the affected VHDX files.
- Host-side WPR traces containing Disk I/O, File I/O, VHDMP, and Hyper-V StorageVSP providers.
- Corresponding guest-side WPR traces.
- A timestamped correlation between guest requests and backing reads.
- Results from a 512e-backed test CSV, if available.
The key question for engineering would be:
Is the low concurrency of decomposed, unaligned reads in VHDMP expected when a 512-byte logical-sector VHDX is hosted on a 4Kn NTFS CSV, or is this a Windows Server 2025 performance issue?
Framing the question this way keeps the focus on the reproducible sector-alignment scenario and avoids assumptions regarding Storage QoS or undocumented StorVSP behavior.
References:
Support policy for 4K sector hard drives - Windows Server | Microsoft Learn
Advanced format (4K) disk compatibility update - Win32 apps | Microsoft Learn
If you find it useful, please click Accept Answer.
Thank you for choosing Microsoft Q&A.
Note: This is the Dutch (nl-NL) forum. I kindly recommend posting your question in Dutch, as this will make it easier for other forum members and subject matter experts to engage with your inquiry. If you prefer to post in English, you may also consider using the Microsoft Q&A English forum, where a larger English-speaking audience can review and respond to your question. Thank you for your understanding