An Azure service that is used to provision Windows and Linux virtual machines.
Azure Virtual Machine with SQL Server: Intermittent Write Stalls on Premium SSD P15 — Request for Analysis and Resolution
Problem description
I am experiencing intermittent multi-second write stalls on my Azure virtual machine running SQL Server, which is hosted on a 256 GiB Premium SSD P15 disk. The issue manifests as normal write latencies of about 6–12 milliseconds, but occasionally, certain write operations take between 24 and over 40 seconds, creating deep I/O queues. This pattern has been observed on both my production VM and an isolated clone, suggesting the problem is not workload-specific. I have already ruled out filesystem minifilter interference, and monitored VM and disk utilization metrics, which remain below documented limits during these stalls. I am seeking assistance to analyze the underlying cause of these delays, particularly to determine whether platform or guest storage components are involved, and to identify potential mitigation steps.
Environment
Azure Virtual Machine running Windows, with SQL Server on a 256 GiB Premium SSD P15 disk, located in UK South, Zone 1.
What I've already tried
I have reproduced the issue on both the production VM and an isolated clone. Ran DiskSPD benchmarking with specific parameters to simulate workload. Verified that no filesystem minifilter drivers are involved, and monitored VM and disk utilization metrics, which stayed below the documented thresholds during the observed stalls. Checked for recent VM or disk configuration changes, such as resizing or snapshots, and confirmed no such modifications occurred around the stall times. Collected Windows event logs (StorPort and StorDiag), performance counters, and storage path evidence during the stalls. No platform events indicative of storage timeouts or resets were found during the incidents.
Current status
The issue persists with no clear platform or configuration anomalies identified. The most recent evidence suggests transient I/O path stalls or queue buildup, but no definitive root cause has been determined. I am requesting detailed analysis of storage logs, event data, and performance metrics around the stall intervals to assist in diagnosing and resolving the problem.