Blacksmith incident
US West storage cluster degradation
Blacksmith experienced a minor incident on September 2, 2026 affecting Incremental Docker Builders (us-west Storage Cluster) and Docker Container Cache (us-west Storage Cluster) and 1 more component, lasting 3h 9m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating Sep 02, 2026, 04:13 PM UTC
We are investigating reports of higher latency on sticky disk operations in our us-west region. Customers running jobs in us-west may see slower incremental Docker builds, Git caching, and container caching, so affected jobs can take longer than usual to complete.
- investigating Sep 02, 2026, 04:48 PM UTC
Sticky disk storage in our us-west region is experiencing higher latency; our other regions are not affected. Customers running jobs in us-west may see slower Git-cached checkouts, incremental Docker builds, and container caching. We are continuing to investigate the underlying cause and will provide another update within the next 30 minutes.
- investigating Sep 02, 2026, 05:22 PM UTC
We are still working to resolve elevated latency on sticky disk storage in our us-west region; other regions are not affected. Customers running jobs in us-west may continue to see slower Git-cached checkouts, incremental Docker builds, and container caching. Our investigation into the underlying cause is ongoing and we will provide another update within the next 30 minutes.
- monitoring Sep 02, 2026, 06:10 PM UTC
Our us-west region storage cluster's latency and error rates have returned to baseline as of approximately 17:15 UTC, following mitigations that reduce load on the affected storage. Git-cached checkouts, incremental Docker builds, and container caching in us-west are back to normal, and we are monitoring to confirm the recovery holds while we continue to investigate the underlying cause. We will provide a final update within the next hour.
- resolved Sep 02, 2026, 07:23 PM UTC
This incident is resolved. Sticky disk storage in our us-west region came under more read load than it could serve at normal latency, which slowed and produced higher error rates for Git-cached checkouts, incremental Docker builds, and container caching. We reduced and redistributed that load, and performance has been normal since approximately 17:15 UTC. Jobs that failed during the incident can be safely re-run.