Blacksmith incident
Degraded job performance, metrics, and log ingestion in us-west
Blacksmith experienced a minor incident on August 10, 2026 affecting Blacksmith Managed Runners (us-west ARM) and Blacksmith Managed Runners (us-west x86) and 1 more component, lasting 1h 26m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating Aug 10, 2026, 07:17 PM UTC
We are experiencing degradation in our metrics and log ingestion services in our us-west region. We are actively investigating the issue.
- investigating Aug 10, 2026, 07:44 PM UTC
We are continuing to investigate degraded metrics and log ingestion in our us-west region. Customers may still see metrics and logs for their jobs appear missing or delayed in the Blacksmith dashboard, while other regions remain unaffected. We will provide another update within the next 30 minutes.
- investigating Aug 10, 2026, 07:45 PM UTC
We are experiencing degraded network performance in our us-west region, affecting metrics and log ingestion as well as job performance. Jobs in us-west that upload artifacts or transfer large amounts of data may run slower than normal and in some cases hit their configured timeouts and fail. Other regions are not affected, and we are actively investigating the issue.
- monitoring Aug 10, 2026, 08:18 PM UTC
Metrics and log ingestion in our us-west region is recovering, and job performance in the region has returned to normal. We are monitoring to confirm the recovery holds and are continuing to investigate the underlying cause. We will provide an update shortly.
- resolved Aug 10, 2026, 08:44 PM UTC
This incident is resolved, with job performance and the ingestion of metrics and logs in our us-west region stable for the past 30 minutes. Jobs that failed or timed out during the incident can be safely re-run.