- Detected by Pingoru
- Sep 13, 2026, 01:11 AM UTC
- Resolved
- Sep 13, 2026, 01:58 AM UTC
- Duration
- 46m
Affected: US East (us-east x86)
Timeline · 3 updates
-
investigating Sep 13, 2026, 01:11 AM UTC
We are currently observing network stalling in our US East region with runner connectivity against AWS's us-east-1 region. Uploads and downloads to and from ECR and S3 may experience occasional stalling. We are investigating.
-
monitoring Sep 13, 2026, 01:40 AM UTC
We have identified a small subset of runners that are exhibiting this network stalling and have isolated them. We're continuing to monitor recovery.
-
resolved Sep 13, 2026, 01:58 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Sep 11, 2026, 05:30 AM UTC
- Resolved
- Sep 11, 2026, 02:59 PM UTC
- Duration
- 9h 29m
Timeline · 3 updates
-
monitoring Sep 11, 2026, 05:30 AM UTC
Since 06:30 UTC Canonical has had three outages affecting [archive.ubuntu.com](http://archive.ubuntu.com), [us.archive.ubuntu.com](http://us.archive.ubuntu.com) and [security.ubuntu.com](http://security.ubuntu.com). Although they have marked these as resolved, we are still seeing intermittent connection failures to those mirrors, which can cause apt-get steps to fail or hang. In the meantime, pointing apt at [azure.archive.ubuntu.com](http://azure.archive.ubuntu.com) instead of [archive.ubuntu.com](http://archive.ubuntu.com) and [security.ubuntu.com](http://security.ubuntu.com) will unblock affected jobs.
-
monitoring Sep 11, 2026, 02:21 PM UTC
We are deploying a mitigation on our side that routes apt around the affected Canonical mirrors. For the time being pointing apt at [azure.archive.ubuntu.com](http://azure.archive.ubuntu.com) in place of [archive.ubuntu.com](http://archive.ubuntu.com) and [security.ubuntu.com](http://security.ubuntu.com) will unblock affected jobs.
-
resolved Sep 11, 2026, 02:59 PM UTC
This incident has been resolved - Canonical's Ubuntu apt mirrors have recovered and apt installs are succeeding again. We are also rolling out an image change to reduce the impact of upstream mirror outages in future.
Read the full incident report →
- Detected by Pingoru
- Sep 09, 2026, 04:46 PM UTC
- Resolved
- Sep 09, 2026, 06:06 PM UTC
- Duration
- 1h 20m
Affected: Incremental Docker Builders (eu-central Storage Cluster)Incremental Docker Builders (us-west Storage Cluster)Docker Container Cache (eu-central Storage Cluster)Docker Container Cache (us-west Storage Cluster)Sticky Disks (eu-central Storage Cluster)Sticky Disks (us-west Storage Cluster)Sticky Disks (eu-west Storage Cluster)Incremental Docker Builders (eu-west Storage Cluster)Docker Container Cache (eu-west Storage Cluster)DashboardActions Cache (US West Cache)Actions Cache (US East Cache)Actions Cache (EU West Cache)Actions Cache (EU Central Cache)Runtime Build Caching (US West Runtime Build Cache)Runtime Build Caching (EU Central Runtime Build Cache)Runtime Build Caching (EU West Runtime Build Cache)
Timeline · 3 updates
-
investigating Sep 09, 2026, 05:14 PM UTC
We are currently experiencing a major outage with both our caching infrastructure and our dashboards. These components are affected across all regions. We are investigating the issue and will provide an update soon.
-
monitoring Sep 09, 2026, 05:21 PM UTC
We are seeing recovery of our caching infrastructure and our dashboard. We will continue to monitor this situation closely.
-
resolved Sep 09, 2026, 06:06 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Sep 08, 2026, 03:00 PM UTC
- Resolved
- Sep 08, 2026, 04:44 PM UTC
- Duration
- 1h 43m
Affected: Github → API Requests
Timeline · 3 updates
-
investigating Sep 08, 2026, 03:00 PM UTC
Since approximately 12:20 UTC, some jobs in our eu-west region have been failing with the GitHub error "The self-hosted runner lost communication with the server"; other regions are unaffected. Affected jobs can be safely re-run while we continue to investigate the underlying cause.
-
monitoring Sep 08, 2026, 03:31 PM UTC
Between approximately 12:20 and 15:15 UTC, some jobs in our eu-west region failed with the GitHub error "The self-hosted runner lost communication with the server"; other regions were unaffected. Job failures in eu-west returned to normal levels as of 15:15 UTC, and any affected jobs can be safely re-run while we continue to investigate the underlying cause.
-
resolved Sep 08, 2026, 04:44 PM UTC
This incident is resolved. Job failures in our eu-west region returned to normal levels at 15:15 UTC and have remained there since; any jobs that failed between approximately 12:20 and 15:15 UTC with the error "The self-hosted runner lost communication with the server" can be safely re-run.
Read the full incident report →
- Detected by Pingoru
- Sep 02, 2026, 04:13 PM UTC
- Resolved
- Sep 02, 2026, 07:23 PM UTC
- Duration
- 3h 9m
Affected: Incremental Docker Builders (us-west Storage Cluster)Docker Container Cache (us-west Storage Cluster)Sticky Disks (us-west Storage Cluster)
Timeline · 5 updates
-
investigating Sep 02, 2026, 04:13 PM UTC
We are investigating reports of higher latency on sticky disk operations in our us-west region. Customers running jobs in us-west may see slower incremental Docker builds, Git caching, and container caching, so affected jobs can take longer than usual to complete.
-
investigating Sep 02, 2026, 04:48 PM UTC
Sticky disk storage in our us-west region is experiencing higher latency; our other regions are not affected. Customers running jobs in us-west may see slower Git-cached checkouts, incremental Docker builds, and container caching. We are continuing to investigate the underlying cause and will provide another update within the next 30 minutes.
-
investigating Sep 02, 2026, 05:22 PM UTC
We are still working to resolve elevated latency on sticky disk storage in our us-west region; other regions are not affected. Customers running jobs in us-west may continue to see slower Git-cached checkouts, incremental Docker builds, and container caching. Our investigation into the underlying cause is ongoing and we will provide another update within the next 30 minutes.
-
monitoring Sep 02, 2026, 06:10 PM UTC
Our us-west region storage cluster's latency and error rates have returned to baseline as of approximately 17:15 UTC, following mitigations that reduce load on the affected storage. Git-cached checkouts, incremental Docker builds, and container caching in us-west are back to normal, and we are monitoring to confirm the recovery holds while we continue to investigate the underlying cause. We will provide a final update within the next hour.
-
resolved Sep 02, 2026, 07:23 PM UTC
This incident is resolved. Sticky disk storage in our us-west region came under more read load than it could serve at normal latency, which slowed and produced higher error rates for Git-cached checkouts, incremental Docker builds, and container caching. We reduced and redistributed that load, and performance has been normal since approximately 17:15 UTC. Jobs that failed during the incident can be safely re-run.
Read the full incident report →
- Detected by Pingoru
- Sep 02, 2026, 02:09 PM UTC
- Resolved
- Sep 02, 2026, 04:30 PM UTC
- Duration
- 2h 21m
Affected: Github → API Requests
Timeline · 4 updates
-
investigating Sep 02, 2026, 02:09 PM UTC
We are currently observing higher rates of errors for the upstream GHCR registries which may indicate an undeclared GitHub incident. We are currently monitoring.
-
investigating Sep 02, 2026, 02:20 PM UTC
We are seeing evidence that Git Checkouts are also affected by this. We are seeing high TCP retransmit rates into the EU GitHub loadbalancer. Checkouts in the EU West, EU Central, US East regions are affected.
-
monitoring Sep 02, 2026, 03:43 PM UTC
Connections to [github.com](http://github.com) and [ghcr.io](http://ghcr.io) from our eu-west, eu-central, and us-east regions have returned to normal as of approximately 15:00 UTC, so checkouts are completing at normal speed and login failures to GitHub Container Registry have dropped to baseline levels. We are continuing to monitor, and re-running any jobs that failed during this period should succeed.
-
resolved Sep 02, 2026, 04:30 PM UTC
This incident has been resolved. GitHub connectivity from our EU regions was degraded from \~12:45 to 16:15 UTC, affecting checkouts, [ghcr.io](http://ghcr.io), and some eu-west job starts. Normal since 16:15\. Affected jobs can be re-run.
Read the full incident report →
- Detected by Pingoru
- Aug 18, 2026, 08:45 PM UTC
- Resolved
- Aug 18, 2026, 11:14 PM UTC
- Duration
- 2h 29m
Timeline · 2 updates
-
monitoring Aug 18, 2026, 08:45 PM UTC
Jobs failure rate at an elevated rate caused by an upstream GitHub being rejected by 429 rate-limit errors. We are seeing a single-digit percentage increase in job failure rate across all organizations and are seeing similar failures on non-Blacksmith infrastructure.
-
resolved Aug 18, 2026, 11:14 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 17, 2026, 06:58 PM UTC
- Resolved
- Aug 17, 2026, 07:00 PM UTC
- Duration
- 1m
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)US Central (us-central MacOS)EU West (eu-west ARM)US East (us-east x86)
Timeline · 2 updates
-
investigating Aug 17, 2026, 06:58 PM UTC
We are seeing an increased error rate processing webhook events which will lead to delays in adoption jobs. We are actively investigating.
-
resolved Aug 17, 2026, 07:00 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 17, 2026, 01:43 PM UTC
- Resolved
- Aug 17, 2026, 08:40 PM UTC
- Duration
- 6h 57m
Affected: Github → API RequestsGithub → WebhooksDashboard
Timeline · 6 updates
-
identified Aug 17, 2026, 01:43 PM UTC
The Blacksmith Dashboard is unable to load. We've identified the root cause to upstream 503s being returned from GitHub on permission-check requests.
-
identified Aug 17, 2026, 01:46 PM UTC
Upstream incident has been declared: We are also seeing elevated error rates in jobs as they hit upstream GitHub errors. Job adoption times are also affected and are delayed. We are monitoring and are looking at potential mitigations.
-
monitoring Aug 17, 2026, 04:56 PM UTC
We're seeing signs of GitHub recovery. The dashboard is now loading and jobs should be running again. We are monitoring the recovery.
-
monitoring Aug 17, 2026, 05:52 PM UTC
We are still seeing intermitting GitHub API errors at a low rate.
-
monitoring Aug 17, 2026, 07:32 PM UTC
We are no longer seeing upstream errors, and are continuing to monitor impact of the upstream outage.
-
resolved Aug 17, 2026, 08:40 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 13, 2026, 11:40 AM UTC
- Resolved
- Aug 14, 2026, 06:14 PM UTC
- Duration
- 1d 6h
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)Actions CacheGithub → WebhooksIncremental Docker Builders (us-west Storage Cluster)Docker Container Cache (us-west Storage Cluster)Sticky Disks (us-west Storage Cluster)EU West (eu-west ARM)US East (us-east x86)Actions Cache (US West Cache)Actions Cache (US East Cache)Actions Cache (EU West Cache)Actions Cache (EU Central Cache)
Timeline · 28 updates
Read the full incident report →
- Detected by Pingoru
- Aug 12, 2026, 06:41 PM UTC
- Resolved
- Aug 13, 2026, 01:20 AM UTC
- Duration
- 6h 38m
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)US Central (us-central MacOS)Github → ActionsEU West (eu-west ARM)
Timeline · 10 updates
Read the full incident report →
- Detected by Pingoru
- Aug 12, 2026, 03:30 AM UTC
- Resolved
- Aug 12, 2026, 03:30 AM UTC
- Duration
- —
Timeline · 1 update
-
resolved Aug 12, 2026, 03:30 AM UTC
Type: Incident Duration: 8 hours and 21 minutes Affected Components: us-west x86 Aug 12, 03:30:00 GMT+0 - Investigating - We are currently investigating this incident. Aug 12, 09:10:10 GMT+0 - Identified - We have identified the affected mirror and are implementing a fix Aug 12, 10:33:01 GMT+0 - Identified - We are currently deploying a mitigation Aug 12, 11:26:10 GMT+0 - Monitoring - We implemented a fix and are seeing improvements. We are continuing to monitor the result. Aug 12, 11:50:47 GMT+0 - Resolved - This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 11, 2026, 07:00 PM UTC
- Resolved
- Aug 11, 2026, 08:00 PM UTC
- Duration
- 59m
Affected: Actions Cache
Timeline · 2 updates
-
investigating Aug 11, 2026, 07:00 PM UTC
We are observing certain customers experiencing higher than baseline cache miss rates.
-
resolved Aug 11, 2026, 08:00 PM UTC
We resolved the configuration error leading to the increased rate of cache misses. Cache hit rates are now at the baseline.
Read the full incident report →
- Detected by Pingoru
- Aug 10, 2026, 07:17 PM UTC
- Resolved
- Aug 10, 2026, 08:44 PM UTC
- Duration
- 1h 26m
Affected: US West (us-west ARM)US West (us-west x86)Dashboard
Timeline · 5 updates
-
investigating Aug 10, 2026, 07:17 PM UTC
We are experiencing degradation in our metrics and log ingestion services in our us-west region. We are actively investigating the issue.
-
investigating Aug 10, 2026, 07:44 PM UTC
We are continuing to investigate degraded metrics and log ingestion in our us-west region. Customers may still see metrics and logs for their jobs appear missing or delayed in the Blacksmith dashboard, while other regions remain unaffected. We will provide another update within the next 30 minutes.
-
investigating Aug 10, 2026, 07:45 PM UTC
We are experiencing degraded network performance in our us-west region, affecting metrics and log ingestion as well as job performance. Jobs in us-west that upload artifacts or transfer large amounts of data may run slower than normal and in some cases hit their configured timeouts and fail. Other regions are not affected, and we are actively investigating the issue.
-
monitoring Aug 10, 2026, 08:18 PM UTC
Metrics and log ingestion in our us-west region is recovering, and job performance in the region has returned to normal. We are monitoring to confirm the recovery holds and are continuing to investigate the underlying cause. We will provide an update shortly.
-
resolved Aug 10, 2026, 08:44 PM UTC
This incident is resolved, with job performance and the ingestion of metrics and logs in our us-west region stable for the past 30 minutes. Jobs that failed or timed out during the incident can be safely re-run.
Read the full incident report →
- Detected by Pingoru
- Aug 06, 2026, 03:30 PM UTC
- Resolved
- Aug 07, 2026, 12:57 AM UTC
- Duration
- 9h 26m
Affected: Github → ActionsGithub → Webhooks
Timeline · 2 updates
-
monitoring Aug 06, 2026, 03:30 PM UTC
Github has reported degraded performance for Actions, jobs may take a moment to be adopted. We are monitoring this incident.
-
resolved Aug 07, 2026, 12:57 AM UTC
GitHub job success and adoption rates are now back at normal levels. Jobs that were not adopted during the incident are being requeued by our team and should be picked up shortly.
Read the full incident report →
- Detected by Pingoru
- Aug 06, 2026, 01:18 AM UTC
- Resolved
- Aug 06, 2026, 02:53 AM UTC
- Duration
- 1h 34m
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)Actions CacheUS Central (us-central MacOS)WebsiteIncremental Docker Builders (eu-central Storage Cluster)Incremental Docker Builders (us-west Storage Cluster)Docker Container Cache (eu-central Storage Cluster)Docker Container Cache (us-west Storage Cluster)Sticky Disks (eu-central Storage Cluster)Sticky Disks (us-west Storage Cluster)Sticky Disks (eu-west Storage Cluster)Incremental Docker Builders (eu-west Storage Cluster)Docker Container Cache (eu-west Storage Cluster)Website (https://blacksmith.sh)CodesmithEU West (eu-west ARM)Runtime Build Caching
Timeline · 6 updates
-
investigating Aug 06, 2026, 01:18 AM UTC
We're currently experiencing an outage of our control plane. GitHub job adoption and execution are affected, as well as dashboard access.
-
identified Aug 06, 2026, 01:36 AM UTC
We've mitigated the root cause and jobs are resuming to run. Some jobs may still be delayed to start as we catch up with the job backlog.
-
monitoring Aug 06, 2026, 01:51 AM UTC
Jobs adoption has recovered and are operational.
-
monitoring Aug 06, 2026, 02:11 AM UTC
There is still a delay in job adoption times as we continue to recover.
-
monitoring Aug 06, 2026, 02:45 AM UTC
Job adoption is now fully operational.
-
resolved Aug 06, 2026, 02:53 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 05, 2026, 10:45 PM UTC
- Resolved
- Aug 05, 2026, 10:45 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Aug 05, 2026, 10:45 PM UTC
Type: Incident Duration: 53 minutes Affected Components: eu-west Storage Cluster, us-west Storage Cluster, , eu-west Storage Cluster, eu-west Storage Cluster, , eu-central Storage Cluster, eu-central Storage Cluster, eu-central Storage Cluster, us-west Storage Cluster, us-west Storage Cluster, Runtime Build Caching → Actions Cache → Aug 5, 22:45:00 GMT+0 - Investigating - We are seeing elevated error rates for Sticky Disk and Actions Cache requests. We are investigating. Aug 5, 22:55:00 GMT+0 - Monitoring - We implemented a fix and are currently monitoring the result. Aug 5, 23:30:42 GMT+0 - Monitoring - We are still monitoring recovery. Aug 5, 23:37:56 GMT+0 - Resolved - This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 03, 2026, 03:37 PM UTC
- Resolved
- Aug 03, 2026, 03:42 PM UTC
- Duration
- 5m
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)US Central (us-central MacOS)EU West (eu-west ARM)
Timeline · 2 updates
-
investigating Aug 03, 2026, 03:37 PM UTC
We're seeing a number of cases where jobs may not be adopted promptly. We're currently investigating.
-
resolved Aug 03, 2026, 03:42 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Aug 02, 2026, 06:16 AM UTC
- Resolved
- Aug 02, 2026, 06:37 AM UTC
- Duration
- 20m
Affected: Incremental Docker Builders (us-west Storage Cluster)Docker Container Cache (us-west Storage Cluster)Sticky Disks (us-west Storage Cluster)
Timeline · 3 updates
-
investigating Aug 02, 2026, 06:16 AM UTC
We are seeing some failures with sticky disk availability and write latency with the US-West storage cluster.
-
monitoring Aug 02, 2026, 06:25 AM UTC
We've applied a fix and are seeing failure rates and latency starting to come down.
-
resolved Aug 02, 2026, 06:37 AM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Jul 30, 2026, 01:30 AM UTC
- Resolved
- Jul 30, 2026, 01:30 AM UTC
- Duration
- —
Timeline · 1 update
-
resolved Jul 30, 2026, 01:30 AM UTC
Type: Maintenance Duration: 1 hour and 44 minutes Affected Components: , Runtime Build Caching → Jul 30, 01:15:54 GMT+0 - Identified - We will be performing maintenance on our Runtime Build Caching infrastructure. Jobs will continue to run, but some may run without remote caching and take longer than usual during this window. No customer action is required. We expect maintenance to last approximately two hours. Jul 30, 01:30:01 GMT+0 - Identified - Maintenance is now in progress Jul 30, 03:13:57 GMT+0 - Completed - Maintenance has completed successfully. Runtime Build Caching has been fully restored and is operating normally.
Read the full incident report →
- Detected by Pingoru
- Jul 29, 2026, 04:05 PM UTC
- Resolved
- Jul 29, 2026, 04:31 PM UTC
- Duration
- 26m
Affected: Codesmith
Timeline · 3 updates
-
identified Jul 29, 2026, 04:05 PM UTC
We're seeing increased error rates from an upstream provider in spinning up sandbox environments in Codesmith. We are looking into remediation options.
-
monitoring Jul 29, 2026, 04:08 PM UTC
The upstream issue is resolved, we are monitoring the sandbox creation rate recovering.
-
resolved Jul 29, 2026, 04:31 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Jul 24, 2026, 07:17 PM UTC
- Resolved
- Jul 24, 2026, 07:51 PM UTC
- Duration
- 33m
Affected: EU Central (eu-central ARM)EU Central (eu-central x86)US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)US Central (us-central MacOS)EU West (eu-west ARM)
Timeline · 4 updates
-
investigating Jul 24, 2026, 07:17 PM UTC
We are seeing webhooks not being accepted by our control plane resulting in jobs not being adopted. We are currently investigating.
-
monitoring Jul 24, 2026, 07:20 PM UTC
We have deployed a fix and are monitoring recovery, and will provide another update within the next 30 minutes.
-
monitoring Jul 24, 2026, 07:23 PM UTC
We have requeued the webhooks that were not processed during the incident, and any affected jobs should start shortly.
-
resolved Jul 24, 2026, 07:51 PM UTC
This incident has been resolved, and job pickup has returned to normal.
Read the full incident report →
- Detected by Pingoru
- Jul 23, 2026, 03:00 PM UTC
- Resolved
- Jul 23, 2026, 03:44 PM UTC
- Duration
- 43m
Affected: Sticky Disks (eu-west Storage Cluster)Docker Container Cache (eu-west Storage Cluster)
Timeline · 3 updates
-
investigating Jul 23, 2026, 03:00 PM UTC
We're seeing large volumes of traffic in our storage cluster in EU-West and this is causing a number of sticky disk related requests to time out and are actively investigating.
-
monitoring Jul 23, 2026, 03:18 PM UTC
We implemented a fix and are currently monitoring the result.
-
resolved Jul 23, 2026, 03:44 PM UTC
This incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Jul 22, 2026, 05:45 PM UTC
- Resolved
- Jul 22, 2026, 06:05 PM UTC
- Duration
- 20m
Affected: US West (us-west ARM)US West (us-west x86)EU West (eu-west x86)
Timeline · 2 updates
-
investigating Jul 22, 2026, 05:45 PM UTC
We are currently experiencing elevated tail latencies for certain customers due to job prioritization. We are noticing jobs are taking upwards of 10m to adopt in a few cases. We are looking at options to alleviate queueing.
-
resolved Jul 22, 2026, 06:05 PM UTC
We have implemented some mitigations and queue times are back to normal. The incident has been resolved.
Read the full incident report →
- Detected by Pingoru
- Jul 22, 2026, 01:00 PM UTC
- Resolved
- Jul 22, 2026, 01:00 PM UTC
- Duration
- —
Timeline · 1 update
-
resolved Jul 22, 2026, 01:00 PM UTC
Type: Incident Duration: 1 hour and 27 minutes Affected Components: , Actions Cache → Jul 22, 13:18:35 GMT+0 - Monitoring - We implemented a fix and are currently monitoring the result. Jul 22, 13:00:00 GMT+0 - Monitoring - We identified an issue which was leading to actions/cache requests experiencing an elevated failure rate. This has now been mitigated and we are monitoring. Jul 22, 14:14:00 GMT+0 - Resolved - This incident has been resolved.
Read the full incident report →