Microsoft incident · via Microsoft Azure

ExpressRoute Gateway - Multiple services experiencing connectivity issues in multiple regions

Major Resolved View vendor source →

Microsoft experienced a major incident on September 30, 2026, lasting 2h 50m. The incident has been resolved; the full update timeline is below.

Started
Sep 30, 2026, 08:30 PM UTC
Resolved
Sep 30, 2026, 11:20 PM UTC
Duration
2h 50m
Detected by Pingoru
Sep 30, 2026, 08:30 PM UTC

Update timeline

  1. monitoring Sep 30, 2026, 08:30 PM UTC

    Starting at 20:30 UTC on 30 September 2026, customers using ExpressRoute and/or Azure VPN Gateway may experience degraded or interrupted connectivity. We are also seeing impact to components responsible for managing network gateways, which may affect some management operations. Our investigation has identified a correlation between the onset of impact and infrastructure ‘operating system’ servicing activity. During this activity, some network gateway instances became unhealthy or temporarily unavailable. We have observed different recovery behavior across the affected services. Many impacted ExpressRoute gateways are recovering as the servicing operation completes. However, some supporting management components have not recovered automatically, and we are actively working to restore those components. Our current working hypothesis is that infrastructure operating system servicing is triggering an unexpected condition in some network service instances. We have not yet confirmed the underlying failure mechanism and are continuing to investigate why affected instances became unhealthy and why recovery behavior differs between services. We have paused further operating system servicing associated with this activity to prevent new instances from entering the affected servicing workflow. In parallel, our engineering teams are restoring network management components that have not recovered automatically as well as evaluating additional recovery actions for resources that remain unhealthy. We are continuing to monitor recovery across the affected services and will provide additional information within the next 30 minutes or as events warrant.