Blacksmith experienced a minor incident on September 2, 2026, lasting —. The incident has been resolved; the full update timeline is below.
Update timeline
- resolved Sep 02, 2026, 04:13 PM UTC
Type: Incident Duration: 3 hours and 10 minutes Affected Components: us-west Storage Cluster, us-west Storage Cluster, us-west Storage Cluster Sep 2, 16:13:47 GMT+0 - Investigating - We are investigating reports of higher latency on sticky disk operations in our us-west region. Customers running jobs in us-west may see slower incremental Docker builds, Git caching, and container caching, so affected jobs can take longer than usual to complete. Sep 2, 16:48:35 GMT+0 - Investigating - Sticky disk storage in our us-west region is experiencing higher latency; our other regions are not affected. Customers running jobs in us-west may see slower Git-cached checkouts, incremental Docker builds, and container caching. We are continuing to investigate the underlying cause and will provide another update within the next 30 minutes. Sep 2, 17:22:52 GMT+0 - Investigating - We are still working to resolve elevated latency on sticky disk storage in our us-west region; other regions are not affected. Customers running jobs in us-west may continue to see slower Git-cached checkouts, incremental Docker builds, and container caching. Our investigation into the underlying cause is ongoing and we will provide another update within the next 30 minutes. Sep 2, 18:10:23 GMT+0 - Monitoring - Our us-west region storage cluster's latency and error rates have returned to baseline as of approximately 17:15 UTC, following mitigations that reduce load on the affected storage. Git-cached checkouts, incremental Docker builds, and container caching in us-west are back to normal, and we are monitoring to confirm the recovery holds while we continue to investigate the underlying cause. We will provide a final update within the next hour. Sep 2, 19:23:31 GMT+0 - Resolved - This incident is resolved. Sticky disk storage in our us-west region came under more read load than it could serve at normal latency, which slowed and produced higher error rates for Git-cached checkouts, incremental Docker builds, and container caching. We reduced and redistributed that load, and performance has been normal since approximately 17:15 UTC. Jobs that failed during the incident can be safely re-run.