Blacksmith incident

Degraded job performance, metrics, and log ingestion in us-west

Minor Resolved View vendor source →

Blacksmith experienced a minor incident on August 10, 2026, lasting —. The incident has been resolved; the full update timeline is below.

Started
Aug 10, 2026, 07:17 PM UTC
Resolved
Aug 10, 2026, 07:17 PM UTC
Duration
Detected by Pingoru
Aug 10, 2026, 07:17 PM UTC

Update timeline

  1. resolved Aug 10, 2026, 07:17 PM UTC

    Type: Incident Duration: 1 hour and 27 minutes Affected Components: us-west x86, us-west ARM, Dashboard Aug 10, 19:17:34 GMT+0 - Investigating - We are experiencing degradation in our metrics and log ingestion services in our us-west region. We are actively investigating the issue. Aug 10, 20:44:19 GMT+0 - Resolved - This incident is resolved, with job performance and the ingestion of metrics and logs in our us-west region stable for the past 30 minutes. Jobs that failed or timed out during the incident can be safely re-run. Aug 10, 19:44:03 GMT+0 - Investigating - We are continuing to investigate degraded metrics and log ingestion in our us-west region. Customers may still see metrics and logs for their jobs appear missing or delayed in the Blacksmith dashboard, while other regions remain unaffected. We will provide another update within the next 30 minutes. Aug 10, 19:45:37 GMT+0 - Investigating - We are experiencing degraded network performance in our us-west region, affecting metrics and log ingestion as well as job performance. Jobs in us-west that upload artifacts or transfer large amounts of data may run slower than normal and in some cases hit their configured timeouts and fail. Other regions are not affected, and we are actively investigating the issue. Aug 10, 20:18:51 GMT+0 - Monitoring - Metrics and log ingestion in our us-west region is recovering, and job performance in the region has returned to normal. We are monitoring to confirm the recovery holds and are continuing to investigate the underlying cause. We will provide an update shortly.