Harness incident

Slowness in Prod1 and Prod2 environment

Minor Resolved View vendor source →

Harness experienced a minor incident on September 8, 2026 affecting Continuous Delivery (CD) - FirstGen - EOS and Continuous Delivery (CD) - FirstGen - EOS and 1 more component, lasting 6h 38m. The incident has been resolved; the full update timeline is below.

Started
Sep 08, 2026, 09:52 PM UTC
Resolved
Sep 09, 2026, 04:30 AM UTC
Duration
6h 38m
Detected by Pingoru
Sep 08, 2026, 09:52 PM UTC

Affected components

Continuous Delivery (CD) - FirstGen - EOSContinuous Delivery (CD) - FirstGen - EOSContinuous Delivery - Next Generation (CDNG)Continuous Delivery - Next Generation (CDNG)Cloud Cost Management (CCM)Cloud Cost Management (CCM)Continuous Error Tracking (CET)Continuous Error Tracking (CET)Chaos EngineeringChaos Engineering

Update timeline

  1. investigating Sep 08, 2026, 09:52 PM UTC

    We are currently investigating this issue.

  2. investigating Sep 08, 2026, 10:53 PM UTC

    We are continuing to investigate this issue.

  3. investigating Sep 08, 2026, 10:53 PM UTC

    We are continuing to investigate the issue.

  4. investigating Sep 08, 2026, 11:31 PM UTC

    We are continuing to investigate this issue.

  5. identified Sep 08, 2026, 11:52 PM UTC

    The issue has been identified and a fix is being implemented.

  6. identified Sep 09, 2026, 12:29 AM UTC

    We are continuing to work on a fix for this issue.

  7. monitoring Sep 09, 2026, 12:36 AM UTC

    A fix has been implemented and we are monitoring the results.

  8. monitoring Sep 09, 2026, 12:39 AM UTC

    A fix has been implemented and we are monitoring the results.

  9. resolved Sep 09, 2026, 04:30 AM UTC

    This incident has been resolved.

  10. postmortem Sep 09, 2026, 11:23 PM UTC

    ### Summary On September 8, 2026, customers in Prod 1 and Prod 2 experienced elevated platform latency and pipeline failures. The issue was caused by a regression in a newly released capability that triggered cascading failures under high load. Because the capability was behind a feature flag, it was quickly disabled, and service was restored after a brief monitoring period. ### Customer Impact * Customers encountered slowness and failures during pipeline execution and UI operations. Some API calls returned errors or timed out. * No data loss or corruption occurred. ### Root Cause The new capability introduced a regression that created contention on a shared backend resource used by multiple Harness components. This saturated the shared platform infrastructure and caused the cascading failures. ### Mitigation * Disabled the capability across all environments * Temporarily increased platform capacity to restore stability ### Next Steps To prevent recurrence, Harness will: 1. **Permanently fix the capability** by profiling and eliminating the sub-optimal code path and query 2. **Improve detection** by enhancing alerting for resource-intensive queries on high-frequency platform paths