Monitoring – Intermittent Service Availability
Timeline · 3 updates
- monitoring Sep 26, 2026, 03:51 PM UTC
At 15:23 UTC we began investigating reports of some services being intermittently unavailable. Service has since been fully restored and all systems are operating normally. We are monitoring closely and continuing to investigate the root cause.
- resolved Sep 26, 2026, 04:49 PM UTC
The incident has been resolved. We are conducting a more thorough investigation to discover the root cause.
- postmortem Sep 28, 2026, 07:03 PM UTC
**Post-incident summary** On September 26, 2026, from 15:28 to 15:39 UTC, requests to our APIs failed sporadically. A large portion of requests were refused outright for several minutes. A few regions briefly recovered partial capacity at 15:30 UTC. What happened: a server in an internal system that our load balancers depend upon ran out of disk space but didn’t fail over properly to a peer node. When that system stopped answering, various load balancers across the regions began restarting automatically. Our engineers moved the affected system to a healthy server at 15:37 UTC, and service was fully restored by 15:39 UTC. What we are doing: we have increased disk capacity and changed log retention on the affected system so it cannot run out of space this way again. We are also changing how our load balancers behave when that system is unavailable, so they keep serving traffic instead of shutting down, and we are adding alerting that would have caught the problem before it affected you.