Major outage: render requests failing
Timeline · 4 updates
- investigating Oct 01, 2026, 09:33 AM UTC
Since 09:33 UTC, most new render requests (screenshots, PDFs and videos) are not completing and are timing out. Cached renders are still served. We have identified the cause in our job-queue infrastructure and are working with our provider to restore it. Webhook deliveries for affected renders will be delayed. We will post another update within 30 minutes.
- identified Oct 01, 2026, 10:12 AM UTC
We are treating this as a major outage. Since 09:33 UTC our job-queue datastore, run by our managed hosting provider, has been accepting connections but not completing them, so new render requests cannot be processed and are timing out. Service has briefly recovered at times and then failed again. The problem is on our provider's side, and we have escalated it to them with high priority. Cached renders are still being served. Queued webhook deliveries will be sent once service is restored. Our next update will be within 30 minutes.
- monitoring Oct 01, 2026, 10:31 AM UTC
Service has recovered. Our provider's job-queue datastore became reachable again, and from 10:11 UTC new render requests have been completing normally across all regions. All of our render monitors were passing again by 10:17 UTC. Jobs and webhook deliveries that were queued during the outage have been processed. Some requests made between 09:33 and 10:11 UTC timed out and may need to be retried. We are monitoring closely and will post a full summary once the incident is resolved.
- resolved Oct 01, 2026, 11:43 AM UTC
Resolved. Render processing has been fully back to normal since 10:11 UTC and has stayed stable since. Between 09:33 and 10:11 UTC (about 38 minutes), the datastore behind our render job queues stopped responding, so new screenshot, PDF and video renders did not complete. Renders and webhooks queued during that window were all processed by 10:42 UTC. Synchronous requests made between 09:33 and 10:11 UTC may have timed out and can be retried. We have since added capacity and redundancy to that datastore. We are working with our provider to confirm the root cause, and will share what we are changing to prevent this happening again.