Blacksmith incident

Delays in job adoption

Major Resolved

Blacksmith experienced a major incident on July 21, 2026 affecting Blacksmith Managed Runners (eu-central ARM) and Blacksmith Managed Runners (eu-central x86) and 1 more component, lasting 9h 5m. The incident has been resolved; the full update timeline is below.

Started
Jul 21, 2026, 02:25 PM UTC
Resolved
Jul 21, 2026, 11:31 PM UTC
Duration
9h 5m
Detected by Pingoru
Jul 21, 2026, 02:25 PM UTC

Affected components

Blacksmith Managed Runners (eu-central ARM)Blacksmith Managed Runners (eu-central x86)Blacksmith Managed Runners (us-west ARM)Blacksmith Managed Runners (us-west x86)Blacksmith Managed Runners (eu-west x86)Actions CacheBlacksmith Managed Runners (us-central MacOS)Incremental Docker Builders (eu-central Storage Cluster)Incremental Docker Builders (us-west Storage Cluster)Docker Container Cache (eu-central Storage Cluster)

Update timeline

  1. investigating Jul 21, 2026, 02:25 PM UTC

    We are currently investigating an issue which is affecting Caching and Monitors for some customers.

  2. identified Jul 21, 2026, 03:44 PM UTC

    We are continuing to investigate the delays in webhook processing on our control plane. Customer's may also see some degradation with caching and sticky disk components.

  3. identified Jul 21, 2026, 05:49 PM UTC

    We're still seeing large scale degradation over the backend. likely most caching + storage agent requests are still failing which is causing to queueing due to the slowdown in builds. Actively investigating and working on mitigations.

  4. investigating Jul 21, 2026, 06:39 PM UTC

    We have rolled out fixes to our workload to bring database query latency back to baseline. The degradation has now moved to other parts of our stack. We are continuing to investigate the root cause here. Customers can still expect to see degraded cache interactions. We will provide an update within the next 30 minutes.

  5. investigating Jul 21, 2026, 07:03 PM UTC

    We have applied a change to our backend systems and metrics are showing partial improvement, though not yet back to baseline. Customers can still expect degraded cache interactions while we continue to investigate the root cause. We will provide an update within the next 30 minutes.

  6. monitoring Jul 21, 2026, 07:38 PM UTC

    Our primary Redis instance, which backs our control plane, hit a saturation point, leading to a feedback loop of load. The initial cause was a burst of deliveries of delayed webhooks from GitHub, and our reconciliation systems added further load to the control plane, making things worse. We have improved load balancing by spreading this workload across multiple Redis instances, and our control plane has fully recovered. The remaining impact is a backlog of queued jobs that we are actively draining, so some customers may still see delayed job starts until the queue clears. We will continue monitoring and update with our findings in the next 30 minutes.

  7. monitoring Jul 21, 2026, 07:56 PM UTC

    Cache and sticky disk operations are healthy again. Bazel caching is not yet operational but being actively investigated by our team. The remaining impact is a backlog of queued jobs that we are actively draining, so some customers may still see delayed job starts until the queue drains.

  8. monitoring Jul 21, 2026, 09:25 PM UTC

    Job in us-west and eu-west have returned to normal. Customers in eu-central may still see delayed job starts while we clear the remaining queue backlog. We have identified the cause of the Bazel caching issue and are now implementing a fix. We will provide another update within the next hour.

  9. monitoring Jul 21, 2026, 10:30 PM UTC

    We are still watching over the job queue draining in one of our regions (eu-central). Customers in this region will see a delay for their jobs to be picked up.

  10. monitoring Jul 21, 2026, 10:59 PM UTC

    We are seeing job adoption return to normal latencies across all regions.

  11. monitoring Jul 21, 2026, 11:19 PM UTC

    We are currently scanning for any missed jobs during this degradation period and ensuring that they are being dispatched.

  12. resolved Jul 21, 2026, 11:31 PM UTC

    This incident has been resolved.