Harness incident

All modules are running slow in Prod1/2/3/4 due to cloud provider incident

Major Resolved View vendor source →

Harness experienced a major incident on August 20, 2026 affecting Platform and Platform and 1 more component, lasting 2h 6m. The incident has been resolved; the full update timeline is below.

Started
Aug 20, 2026, 05:17 PM UTC
Resolved
Aug 20, 2026, 07:23 PM UTC
Duration
2h 6m
Detected by Pingoru
Aug 20, 2026, 05:17 PM UTC

Affected components

PlatformPlatformPlatformPlatform

Update timeline

  1. investigating Aug 20, 2026, 03:37 PM UTC

    We are currently investigating this issue.

  2. investigating Aug 20, 2026, 04:05 PM UTC

    The slowness could cause either of below symptoms: - Pipelines not starting - Delays in execution - Pipelines being cancelled due to timeouts

  3. investigating Aug 20, 2026, 04:22 PM UTC

    Our cloud provider is facing an active incident and we are following up.

  4. investigating Aug 20, 2026, 04:32 PM UTC

    We are continuing to investigate this issue.

  5. identified Aug 20, 2026, 04:50 PM UTC

    Our cloud provider has confirmed an ongoing incident impacting multiple regions. Harness pipelines have not experienced failures as a result, though some users may continue to experience slowness. We are monitoring the situation closely and will provide updates as more information becomes available.

  6. monitoring Aug 20, 2026, 05:17 PM UTC

    We are observing better latencies across the board due to the fix at the cloud provider end. We are keeping a close watch.

  7. resolved Aug 20, 2026, 07:23 PM UTC

    This incident has been resolved.

  8. postmortem Aug 26, 2026, 04:57 AM UTC

    # Summary On 20 August 2026, beginning at approximately 15:00 UTC, the Harness platform experienced widespread performance degradation across all production environments. Pipeline executions that normally complete in around two minutes took seven to ten minutes. Continuous Delivery, Continuous Integration, pipeline orchestration, and Feature Management & Experimentation were all affected. Google Cloud Platform experienced a multi-product incident in the us-west1 region affecting Bigtable, Compute Engine, Google Kubernetes Engine, and persistent-disk I/O. Harness production infrastructure runs on persistent disks in that region. The degradation raised database operation latency from approximately 2 ms to over 10 ms at the 95th percentile, which in turn caused message-queue processing lag and propagated to every service that depends on timely database access. ‌ # Impact This was a degradation, not an outage. Pipelines continued to execute and complete successfully throughout; they were slow rather than failing. No data was lost, and no customer work was dropped as a result of this incident. # **Root cause** Harness production infrastructure in the affected environments runs on Google Cloud Platform persistent disks in the us-west1 region. When that storage layer degraded, the effect propagated through the platform in a predictable chain: **Persistent-disk I/O degradation in us-west1.** Google Cloud Platform experienced a multi-product incident affecting Bigtable, Compute Engine, Google Kubernetes Engine, and persistent-disk performance. This was an infrastructure failure in the provider’s environment, outside Harness’s control. # **Preventive actions** Although Harness cannot prevent a cloud provider infrastructure failure. The actions below are aimed at detecting one faster and being better positioned to act on it. | **Action** | | --- | | Continue routine pre-testing of targeted cross-region database failovers, as performed during this incident, to keep failover readiness verified rather than assumed | | Assess full-stack multi-region failover readiness for future scenarios in which cross-region latency would be unacceptable |