Harness incident

Few customers experiencing intermittent errors during pipeline execution.

Minor Resolved View vendor source →

Harness experienced a minor incident on July 9, 2026 affecting Continuous Delivery (CD) - FirstGen - EOS and Continuous Integration Enterprise(CIE) - Linux Cloud Builds, lasting 1h 24m. The incident has been resolved; the full update timeline is below.

Started
Jul 09, 2026, 08:21 PM UTC
Resolved
Jul 09, 2026, 09:46 PM UTC
Duration
1h 24m
Detected by Pingoru
Jul 09, 2026, 08:21 PM UTC

Affected components

Continuous Delivery (CD) - FirstGen - EOSContinuous Integration Enterprise(CIE) - Linux Cloud Builds

Update timeline

  1. investigating Jul 09, 2026, 08:21 PM UTC

    We are currently investigating this issue.

  2. monitoring Jul 09, 2026, 09:10 PM UTC

    A fix has been implemented and we are monitoring the results.

  3. resolved Jul 09, 2026, 09:46 PM UTC

    This incident has been resolved.

  4. postmortem Jul 21, 2026, 05:21 AM UTC

    # Summary On July 9, 2026, some customers in the Harness Prod3 cluster experienced pipeline failures over a ~2.5-hour window, despite making no changes in their pipelines. Shell Script steps failed when fetching scripts from the Harness File Store, and Kubernetes deployment steps failed when retrieving secret encryption details. # Impact 1. **Affected users**: Some customers using the Harness Prod3 cluster reported pipeline execution failures. 2. **User impact:** Customers experienced failures in Shell Script and Kubernetes deployment steps. Impacted pipelines required manual retries after the incident was resolved. 3. **Scope:** This was not a platform-wide outage. The issue was isolated to specific APIs in our internal service within the Prod3 cluster. # Root Cause The incident was caused by degraded performance in a downstream service. This led to a storm of timeout errors in our internal service, which caused these failures. # Mitigation Engineering identified the change that caused this issue in the downstream dependency service and deployed a patch to restore it. Once deployed, pipeline executions returned to normal success rates. # Preventive Measures & Next Steps 1. Added additional monitoring and alerting on the downstream service to detect performance degradation earlier. 2. Reviewing timeout and retry logic in the internal service to improve resilience against downstream latency spikes. 3. Evaluating circuit breaker patterns to prevent cascading failures from downstream dependencies.