Dstny incident

Call2Teams - West US region intermittent networking outage

Minor Resolved View vendor source →

Dstny experienced a minor incident on July 23, 2026 affecting US, lasting 2h 50m. The incident has been resolved; the full update timeline is below.

Started
Jul 23, 2026, 05:09 PM UTC
Resolved
Jul 23, 2026, 08:00 PM UTC
Duration
2h 50m
Detected by Pingoru
Jul 23, 2026, 05:09 PM UTC

Affected components

US

Update timeline

  1. investigating Jul 23, 2026, 05:09 PM UTC

    We are investigating a networking issue affecting connectivity to Azure services in the West US region. Impacted customers may experience intermittent connectivity failures, increased latency, or difficulty accessing Azure services. Customers with traffic traversing the West US region may also experience downstream impact.

  2. identified Jul 23, 2026, 07:19 PM UTC

    Microsoft: We identified a recent change that was strongly correlated with the onset of impact. We have completed the rollback of this change and telemetry across services are continuing to show signs of recovery. Customer should be observing recovery at this time. We are closely monitoring service health and downstream recovery and will provide additional updates as we validate restoration This message was last updated at 18:47 UTC on 23 July 2026

  3. resolved Jul 24, 2026, 08:05 AM UTC

    On July 23, 2026, a routine device maintenance required isolating specific network paths. Our maintenance process converts these requests into system-readable requests and verifies that at least one of the two redundant paths remains healthy. Before maintenance begins, these requests undergo safety checks to confirm the work will be impact-less. In this case, a bug in the request conversion system incorrectly marked additional devices as a part of the maintenance event and caused a set of IP routes to be removed from more devices than intended. The routes were removed between our datacenter and wide-area network, impacting traffic entering or exiting the region. At 14:44 UTC, customers started experiencing impact and the engineering team was engaged immediately. The issue initially presented as large-scale route churn in our Wide-Area Network (WAN). It was later found that route removal was from a datacenter in the West US region. Once confirmed, engineers began roll back at 17:45 UTC. By 18:26 UTC, the network was restored and healthy, and all impacted services had fully recovered by 19:41 UTC.

  4. postmortem Jul 30, 2026, 03:41 PM UTC

    **Incident Summary** From 13:49 UTC on 23rd July until 17:37 UTC on 28th July 2026, a subset of users in the US West region experienced intermittent issues registering to Edge SBCs, which impacted the ability to place or receive calls. The issue was proactively detected by our internal monitoring at 14:49 UTC on 23rd July, which identified connectivity failures to specific hosts. Investigation confirmed that the root cause was a regional network outage within the Azure infrastructure. While the primary service impact was mitigated by 18:37 UTC on 23rd July following a vendor-initiated rollback of network changes, the incident remained under observation until final verification was completed on 28th July, with all services verified as stable. **Root Cause** The incident was caused by an external third-party dependency failure. A regional network outage within Microsoft Azure’s US West infrastructure prevented registration traffic from reaching the Edge SBC hosts. This connectivity gap was the result of a network change implemented by the vendor within their environment, which inadvertently disrupted access to multiple hosts in that region. **Incident Resolution** The issue was resolved following a rollback of the problematic network changes by Azure engineering teams. Dstny engineers monitored the restoration of connectivity and confirmed that SSH and registration traffic returned to normal parameters. Throughout the event, the Call2Teams high-availability architecture allowed many customers to transparently fail over to secondary Edge SBCs, significantly limiting the total number of affected users. Final stability checks were performed across the affected region to ensure full service restoration. **Mitigative Actions** Our monitoring and alerting on the SBC platform functioned as intended and enabled the issue affecting our service to be identified and responded to promptly. As part of our continuous improvement activities, we will review opportunities to further enhance visibility of issues affecting external service providers and dependencies, recognising that some vendor-side infrastructure events may not currently be detectable through our existing monitoring capabilities. ### **Timeline** * 23 Jul 2026, 13:49 - Estimated start of service impact. * 23 Jul 2026, 14:49 - Call test alert triggered for US West host; engineering began investigation. * 23 Jul 2026, 15:12 - Affected host taken out of service and restarted to attempt recovery. * 23 Jul 2026, 16:30 - Engineering confirmed multiple hosts in the region were affected by a broader Azure network issue. * 23 Jul 2026, 16:40 - Official vendor notification received regarding US West connectivity issues. * 23 Jul 2026, 17:45 - Vendor initiated a rollback of recent network changes. * 23 Jul 2026, 18:37 - Call test alerts cleared and primary service impact concluded. * 28 Jul 2026, 17:37 - Incident officially closed following extended monitoring and verification.