Harness incident

Intermittent External Network Connectivity Issues Affecting Build VMs

Minor Resolved View vendor source →

Harness experienced a minor incident on July 30, 2026 affecting Continuous Integration Enterprise(CIE) - Linux Cloud Builds and Continuous Integration Enterprise(CIE) - Linux Cloud Builds and 1 more component, lasting 21h 23m. The incident has been resolved; the full update timeline is below.

Started
Jul 30, 2026, 05:59 AM UTC
Resolved
Jul 31, 2026, 03:22 AM UTC
Duration
21h 23m
Detected by Pingoru
Jul 30, 2026, 05:59 AM UTC

Affected components

Continuous Integration Enterprise(CIE) - Linux Cloud BuildsContinuous Integration Enterprise(CIE) - Linux Cloud BuildsContinuous Integration Enterprise(CIE) - Linux Cloud BuildsContinuous Integration Enterprise(CIE) - Linux Cloud BuildsContinuous Integration Enterprise(CIE) - Linux Cloud Builds

Update timeline

  1. investigating Jul 30, 2026, 05:59 AM UTC

    Summary - We are intermittently facing network connectivity issues with our Build VM's unable to connect to external resources. We are currently investigating the issue.

  2. monitoring Jul 30, 2026, 08:07 AM UTC

    A fix has been implemented and we are monitoring the results.

  3. monitoring Jul 30, 2026, 08:07 AM UTC

    We are continuing to monitor for any further issues.

  4. resolved Jul 31, 2026, 03:22 AM UTC

    This incident has been resolved.

  5. postmortem Aug 07, 2026, 07:32 PM UTC

    ## Summary Starting on August 4, 2026, CI runners in the us-west1 and us-central1 regions intermittently experienced connection timeouts of approximately 134 seconds when reaching external services such as GitHub and Bitbucket over outbound network gateways. ## Impact * CI runners in the affected regions intermittently experienced connection timeouts of approximately 134 seconds when reaching external services \(e.g., GitHub, Bitbucket\) over our outbound network gateways. * The issue was intermittent rather than constant — connections succeeded under normal load, and failures clustered during periods of high outbound traffic volume. * No data was lost or corrupted. This was a network-connectivity and capacity issue, not a data-integrity issue. * us-west1 and us-central1 were the affected regions; other regions were not impacted by this issue. ## Root Cause ‌ Our load balancer distributes outbound traffic across multiple NAT gateways using a hashing method based on connection details \(source/destination address and port\). For any single connection, these details stay constant for that connection's lifetime. We had a sustainted traffic surge for a few seconds which congested the gateways ‌ ## Action Items To prevent such issues from happening again Harness will, Increase outbound connection capacity on our NAT gateways by provisioning additional external network interfaces, giving each gateway a substantially larger pool of connections it can serve concurrently..