Microsoft's Azure status history posted two post-incident reviews in quick succession at the end of September. Neither was a global outage, but both root causes are textbook cases worth checking against your own architecture.

Incident One: AI Services in Sweden Central (Tracking ID 7QL5-Z50)

From 10:03 to 15:58 UTC on September 29, 2026, Azure OpenAI, Foundry Agent Service, Foundry Models and Cognitive Services in Sweden Central saw intermittent request failures, increased latency and HTTP 5XX errors. Microsoft's stated root cause: a backend metadata retrieval service hit database timeouts, and its instances reached their thresholds and restarted repeatedly, shrinking healthy capacity. Failure rates were detected at 10:14, the timeouts identified at 11:29, services scaled and resource limits raised at 15:22, and recovery reached at 15:58.

It is the same class of risk we described in our piece on the September 3 AI service disruption: your application is fine, but the model API it depends on is slow or failing in one region.

Incident Two: Gateway Services in Five Regions (Tracking ID 7Q30-010)

From 20:30 UTC on September 30 to 02:15 UTC on October 1, ExpressRoute Gateway, Azure Firewall, Application Gateway, WAF, VPN Gateway and Azure VMware Solution had connectivity problems in UK South, France Central, North Europe, Southeast Asia and UK West — Southeast Asia being Azure's Singapore region. The preliminary root cause: a recent change to regional gateway management triggered excessive load while unrelated OS servicing maintenance was under way, and the service could not auto-scale because of constraints in a dependent service. Microsoft correlated the incident with the OS servicing and paused it at 23:05, recovery progressed at 01:36, and mitigation was confirmed at 02:15.

Reading the Two Together

  • "The machine is up" is not "the service is reachable". In the second incident many virtual machines may have been perfectly healthy, yet with the gateway, firewall or VPN in front of them down, users still couldn't connect. Host-level monitoring isn't enough; probe the full access path from outside.
  • Changes colliding with maintenance is an old failure mode. Either one alone is harmless; together they break things. The same applies to your own servers: stagger upgrades, reboots and config changes, and keep a rollback path.
  • Draw your dependencies. In our report on the first half's outages we suggested mapping which requests pass through which provider's region, and whether you can switch when one fails.

If You Serve Users in Southeast Asia

If your users are mainly in Southeast Asia and your critical entry point sits in a single region of a single provider, an incident like this leaves you waiting. A sturdier setup keeps an entry point that can take over at another provider or in another region, paired with DNS failover. Choosing a low-latency data centre for your workload covers how to pick locations, and our Singapore node can serve as one such standby entry point.