Blacksmith incident

Degraded job performance, metrics, and log ingestion in us-west

Minor Resolved

Blacksmith experienced a minor incident on August 10, 2026 affecting Blacksmith Managed Runners (us-west ARM) and Blacksmith Managed Runners (us-west x86) and 1 more component, lasting 1h 26m. The incident has been resolved; the full update timeline is below.

Started
Aug 10, 2026, 07:17 PM UTC
Resolved
Aug 10, 2026, 08:44 PM UTC
Duration
1h 26m
Detected by Pingoru
Aug 10, 2026, 07:17 PM UTC

Affected components

Blacksmith Managed Runners (us-west ARM)Blacksmith Managed Runners (us-west x86)Observability

Update timeline

  1. investigating Aug 10, 2026, 07:17 PM UTC

    We are experiencing degradation in our metrics and log ingestion services in our us-west region. We are actively investigating the issue.

  2. investigating Aug 10, 2026, 07:44 PM UTC

    We are continuing to investigate degraded metrics and log ingestion in our us-west region. Customers may still see metrics and logs for their jobs appear missing or delayed in the Blacksmith dashboard, while other regions remain unaffected. We will provide another update within the next 30 minutes.

  3. investigating Aug 10, 2026, 07:45 PM UTC

    We are experiencing degraded network performance in our us-west region, affecting metrics and log ingestion as well as job performance. Jobs in us-west that upload artifacts or transfer large amounts of data may run slower than normal and in some cases hit their configured timeouts and fail. Other regions are not affected, and we are actively investigating the issue.

  4. monitoring Aug 10, 2026, 08:18 PM UTC

    Metrics and log ingestion in our us-west region is recovering, and job performance in the region has returned to normal. We are monitoring to confirm the recovery holds and are continuing to investigate the underlying cause. We will provide an update shortly.

  5. resolved Aug 10, 2026, 08:44 PM UTC

    This incident is resolved, with job performance and the ingestion of metrics and logs in our us-west region stable for the past 30 minutes. Jobs that failed or timed out during the incident can be safely re-run.