Intermittent HTTP 500 Internal Server Errors on ARM Deployments (germanywestcentral)

Laurens Bremers 25 Reputation points
2026-03-10T15:59:56.37+00:00

We experienced intermittent Internal Server Errors when executing deployments via Azure Resource Manager on Feb 26 (09:30-16:00 UTC). This issue persisted for several hours with retries resulting in a mix of successes and failures, after which the issue disappeared. We noticed these errors not only in our own Azure Tenant, but also in tenants of our customers that use our deployment scripts to automate their Azure Virtual Desktop deployments (through the Azure Go SDK).

We had a very similar issue on December 18th 2025: intermittent Internal Server Errors during deployments for a few hours, issue being resolved around 16:00 UTC.

During this windows, the Azure Status page reported all services as healthy in this region. Because we do not have a paid support plan, we are unable to open a technical support ticket through the portal to get this platform issue investigated.

Would it be possible for a one-time support ticket to be opened so the root cause can be identified? We suspect a backend node failure or capacity issue in the region.

Some of the event log entries of failed deployments:<REDACTED RESOURCE DATA> <moved to Private Message>

Azure Virtual Machines
Azure Virtual Machines

An Azure service that is used to provision Windows and Linux virtual machines.


Answer accepted by question author
Anonymous
2026-03-11T17:08:31.42+00:00

Hello Laurens,

We appreciate your patience as we worked with the backend team to investigate the root cause of these deployment failures. 

The investigation revealed that on 2026-03-17, there was a transient fault in one backend instance of Azure Resource Manager (ARM) caused failures for any CRUD operations routed to it, while other instances remained unaffected. The issue was detected and resolved automatically after the instance restarted. No configuration changes or deployments triggered the problem, and it was confirmed to be an isolated reliability fault.

To prevent recurrence, the platform team has filed repair items and is working on a fix, aiming for global deployment by May 2026. We apologize for any inconvenience caused.

 

Recommendations:

Implement retry logic to reduce the likelihood of this issue affecting your operations.

Evaluate your application's reliability using guidance from the Azure Well-Architected Framework and its interactive Well-Architected Review: https://aka.ms/AzPIR/WAF

The impact times listed reflect the full duration of the incident and are not specific to individual customers. For advice on monitoring to understand detailed impact: https://aka.ms/AzPIR/Monitoring

Was this answer helpful?

1 person found this answer helpful.

0 additional answers

Sort by: Most helpful

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.