Blacksmith incident
Backend Degradation causing job slowness and caching failures
Blacksmith experienced a minor incident on September 29, 2026 affecting EU Central (eu-central ARM) and EU Central (eu-central x86) and 1 more component, lasting 1h 54m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating Sep 29, 2026, 09:10 PM UTC
Jobs across all regions are taking longer to start and cache operations are failing. We are continuing to investigate the root cause.
- identified Sep 29, 2026, 09:28 PM UTC
We have applied a fix and are monitoring its effect. Customers may still see delayed job starts and failing cache operations while recovery completes. We will provide an update within the next 30 minutes.
- monitoring Sep 29, 2026, 09:43 PM UTC
Our mitigation has taken effect and job start times and cache operations are improving. Customers may still see some delayed job starts and cache failures while recovery completes. We will provide an update within the next 30 minutes.
- monitoring Sep 29, 2026, 10:20 PM UTC
Caching has recovered, and job queue times across EU regions have recovered. Queue times for larger jobs (16 and 32 vcpu jobs) in US west and US east continue to remain elevated. We are continuing to monitor for any regressions.
- monitoring Sep 29, 2026, 10:42 PM UTC
This incident is resolved. Job starts and cache operations returned to normal in all regions, and we are continuing to monitor. Jobs that failed during the incident can be re-run.
- resolved Sep 29, 2026, 11:04 PM UTC
This incident has been resolved.