Plataforma e infraestructura de informática en la nube para crear, implementar y administrar aplicaciones y servicios a través de una red mundial de centros de datos administrados por Microsoft.
Databricks: Clusters Fail During VM Provisioning with InternalExecutionError
Problem description
I am experiencing an issue where multiple production Azure Databricks classic clusters are failing to provision VMs during startup. The failure occurs with an error code 'InternalExecutionError' and a message indicating an internal execution error occurred, advising to retry later. The Azure activity log shows that the VM creation fails with 'ResourceOperationFailure' and nested 'InternalExecutionError'. This problem started on September 30, 2026, and affects clusters requesting on-demand VMs during startup. The environment involves Azure Databricks using classic compute in a production setting, and the failure occurs during VM provisioning. I have not made recent configuration changes or infrastructure modifications that could cause this.
Environment
Azure Databricks workspace using classic compute (clusters), in a production environment.
What I've already tried
I have attempted to manually start clusters without any configuration changes, and I have reviewed the Azure activity logs and Databricks termination reasons. No other troubleshooting steps or retries have been documented in the case.
Current status
The issue persists with multiple clusters failing to start due to VM provisioning errors. I am seeking assistance to investigate the root cause, particularly to determine if this is related to Azure capacity, regional quotas, or platform-side provisioning issues, and to identify potential mitigation steps.