An Azure service that is used to provision Windows and Linux virtual machines.
Azure Support case 2610010050004147 — no RCA or substantive update after 5 days
We experienced an unexpected Azure subscription cancellation followed by an immediate reactivation while 9 virtual machines were running continuous model-training workloads.
According to the Azure Activity Log:
- 12:25:29 PM BST — subscription cancelled
- 12:28:01 PM BST — subscription reactivated
The actions were not initiated by us. As a result, all 9 VMs were stopped with the VMStoppedToWarnSubscription error, and approximately 24 hours of model-training progress was lost.
We opened Azure Support case 2610010050004147 and explicitly requested:
- a Root Cause Analysis explaining why the subscription was cancelled and reactivated;
- identification of the source/initiator of the subscription actions;
- an explanation of why the VMs were stopped;
- preventive measures to ensure this cannot happen again;
- information about possible service credits for the wasted compute time.
The assigned support engineer acknowledged the impact and stated that Microsoft was checking the issue internally with the relevant team. However, after several days and multiple follow-ups, we still have not received any substantive information about the investigation, the team handling it, the root cause, or even a meaningful status update.
My question is:
What is the proper escalation path for an Azure subscription incident when the assigned support engineer confirms that an internal investigation is underway but provides no RCA or substantive progress after several days?
This is not a general technical question. The incident already occurred and caused a significant interruption to an active production workload.
I would especially appreciate an answer from an Azure Support engineer or Microsoft employee regarding:
- How this type of incident should be escalated.
- Whether the customer can request escalation to a support manager or the relevant subscription engineering team.
- What information should be provided to help identify the source of a
Cancel subscription/Reactivate subscriptionoperation. - Whether there is a standard process for reviewing service-credit requests after an Azure-side incident causes significant wasted compute.We experienced an unexpected Azure subscription cancellation followed by an immediate reactivation while 9 virtual machines were running continuous model-training workloads. According to the Azure Activity Log:
- 12:25:29 PM BST — subscription cancelled
- 12:28:01 PM BST — subscription reactivated
VMStoppedToWarnSubscriptionerror, and approximately 24 hours of model-training progress was lost. We opened Azure Support case 2610010050004147 and explicitly requested:- a Root Cause Analysis explaining why the subscription was cancelled and reactivated;
- identification of the source/initiator of the subscription actions;
- an explanation of why the VMs were stopped;
- preventive measures to ensure this cannot happen again;
- information about possible service credits for the wasted compute time.
- How this type of incident should be escalated.
- Whether the customer can request escalation to a support manager or the relevant subscription engineering team.
- What information should be provided to help identify the source of a
Cancel subscription/Reactivate subscriptionoperation. - Whether there is a standard process for reviewing service-credit requests after an Azure-side incident causes significant wasted compute.