URGENT: Azure OpenAI Services interrupted and mostly not available in Sweden Central

Anonym
2026-01-28T11:01:14.5666667+00:00

Just like yesterday, January 27, at around 09:00 UTC, today, on January 28 at around 09:00 UTC, all communications with GPT-4.1 and GPT-5.x models via the Responses-API (/v1) is interrupted!

So far, Microsoft hasn't recgonized this as a severe service interruption on the Azure Service Health website. Can someone at Microsoft please escalate this?

Azure OpenAI in Foundry-Modellen
Azure OpenAI in Foundry-Modellen

Ein Azure-Dienst, der Zugriff auf die GPT-3-Modelle von OpenAI ermöglicht und Unternehmensfunktionen bietet


Antwort, die vom Frageautor angenommen wurde
SRILAKSHMI C 19,735 Zuverlässigkeitspunkte Externe Microsoft-Mitarbeiter Moderator
2026-01-28T14:20:42.4466667+00:00

Hello Daniel Musil,

Welcome to Microsoft Q&A and Thank you for reaching out.

As per Team, what happened?

Between 09:22 UTC and 16:12 UTC on 27 January 2026, a platform issue resulted in an impact to the Azure OpenAI Service in Sweden Central region. Impacted customers may have seen HTTP 500/503 errors, failed inference requests, and issues with model deployment metadata. This issue also affected downstream AI Services dependent on Azure OpenAI in this region.

What do we know so far?

Our initial investigation indicates that the issue may be related to elevated error handling within one of our production model dependencies. This temporarily affected request processing, specifically authorization of the incoming requests, causing intermittent service degradation impacting request success rates. We mitigated the issue by stabilizing traffic flow and adjustments to improve request handling and resilience, validating system health, and monitoring recovery to ensure normal operation has been restored.

How did we respond?

09:22 UTC on 27 January 2026 – The issue was detected through service monitoring which is also when customers began to see intermittent availability issues were observed.

12:36 UTC on 27 January 2026 – Initiated mitigation to restart the IRM service on the Sweden Central clusters.

12:46 UTC on 27 January 2026 – Identified that Sweden Central cluster is seeing pods crashing with out-of-memory errors.

13:02 UTC on 27 January 2026 – Initiated mitigation workflow by scaling out nodes in the cluster to improve request handling and resilience.

15:30 UTC on 27 January 2026 – Started to increase the memory available in the pods to alleviate memory load on the cluster.

15:53 UTC on 27 January 2026 – Completed increase in memory in the pods to alleviate memory load on the cluster.

16:12 UTC on 27 January 2026 – Service(s) restored, and customer impact mitigated.

What happens next?

  • Our team will be completing an internal retrospective to understand the incident in more detail. Once that is completed, generally within 14 days, we will publish a Post Incident Review (PIR) to all impacted customers.

To get notified if a PIR is published, and/or to stay informed about future Azure service issues, make sure that you configure and maintain Azure Service Health alerts – these can trigger emails, SMS, push notifications, webhooks, and more: https://aka.ms/ash-alerts For more information on Post Incident Reviews, refer to https://aka.ms/AzurePIRs The impact times above represent the full incident duration, so are not specific to any individual customer. Actual impact to service availability may vary between customers and resources – for guidance on implementing monitoring to understand granular impact: https://aka.ms/AzPIR/Monitoring

Thank you!

War diese Antwort hilfreich?

0 Kommentare Keine Kommentare

1 zusätzliche Antwort

Sortieren nach: Am hilfreichsten
  1. Anonym
    2026-01-29T12:30:17.6166667+00:00

    It's January 29. This is the THIRD day in a row on which GPT-4.1 and GPT-5.x models in Sweden Central have service interruptions. None of these issues even appear on the Azure Service Health sites. How is this possible!?

    War diese Antwort hilfreich?

    0 Kommentare Keine Kommentare

Ihre Antwort

Antworten können von Fragestellenden als „Angenommen“ und von Moderierenden als „Empfohlen“ gekennzeichnet werden, wodurch Benutzende wissen, dass diese Antwort das Problem des Fragestellenden gelöst hat.