Twingate experienced a critical incident on May 28, 2026 affecting Authentication - Enterprise and Authentication - Social and 1 more component, lasting 8h 1m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating May 28, 2026, 09:24 AM UTC
We are currently investigating this issue.
- investigating May 28, 2026, 09:25 AM UTC
We are continuing to investigate this issue.
- investigating May 28, 2026, 10:55 AM UTC
We are continuing to investigate this issue.
- investigating May 28, 2026, 10:57 AM UTC
We are continuing to investigate this issue.
- monitoring May 28, 2026, 12:02 PM UTC
The issue has been fixed. We are still working with cloud provider on root cause.
- monitoring May 28, 2026, 01:47 PM UTC
We're continuing to work with our cloud provider to identify the root cause.
- monitoring May 28, 2026, 04:22 PM UTC
We're continuing to work with our cloud provider to identify the root cause.
- resolved May 28, 2026, 05:27 PM UTC
We are currently awaiting the completion of our cloud provider’s investigation into the issue. In the meantime, we continue to closely monitor the system, and all service metrics have returned to normal operating levels. At this time, we are closing this incident. We will publish a comprehensive incident report as soon as it becomes available.
- postmortem Jun 03, 2026, 04:39 PM UTC
# Incident Report – Authorization Service Degradation ## Components Impacted Control Plane – Authorization Service ## Summary On May 28, 2026, between 10:00 UTC and 12:30 UTC, Twingate experienced a degradation of its Authorization service that affected approximately 15% of active connections at peak impact. During the incident, authorization requests experienced elevated network latency. As request processing times increased, the Authorization service's effective capacity to handle incoming traffic was reduced, resulting in elevated error rates and intermittent authorization failures for a subset of customers. Twingate operates the Authorization service across multiple cloud regions in an active-active configuration. While the platform remained available throughout the event, the combination of increased request latency and reduced service capacity led to customer impact until additional capacity was provisioned and service performance stabilized. ## Root Cause The incident was triggered by elevated network latency affecting communication paths used by the Authorization service. As requests took longer to complete, individual service instances were able to process fewer requests than normal. This reduction in throughput exposed a limitation in our auto-scaling configuration, which primarily relied on CPU utilization to determine service capacity requirements. As request-processing workers spent more time waiting on network operations, CPU utilization declined even as request latency increased. As a result, the service scaled down during a period of elevated request latency, reducing available capacity and amplifying customer impact. Recovery efforts were further complicated by an unusually high rate of spot instance preemptions in two regions, which reduced available compute capacity during stabilization. ## Resolution Engineering teams mitigated the incident by manually increasing Authorization service capacity and expanding available cluster resources. As additional capacity came online, request latency and error rates returned to normal levels and service performance fully recovered. ## Corrective Actions ### Completed * Increased the baseline capacity of the Authorization service by raising the minimum number of service instances. * Increased baseline node capacity across affected Kubernetes clusters. * Implemented additional rate controls for unusually large authorization requests to better protect overall service availability during periods of elevated load. ### In Progress * Rebalance spot and on-demand node capacity to reduce sensitivity to spot instance interruptions. * Enhance auto-scaling policies to incorporate latency and service performance metrics in addition to CPU utilization. * Expand monitoring and alerting to better detect conditions where request latency increases while resource utilization decreases.