Harness experienced a minor incident on September 22, 2026 affecting Continuous Integration Enterprise(CIE) - Linux Cloud Builds and Continuous Integration Enterprise(CIE) - Linux Cloud Builds and 1 more component, lasting 1h 57m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating Sep 22, 2026, 04:40 PM UTC
We are currently investigating this issue.
- identified Sep 22, 2026, 05:14 PM UTC
The issue has been identified and a fix is being implemented.
- monitoring Sep 22, 2026, 06:01 PM UTC
A fix has been implemented and we are monitoring the results.
- resolved Sep 22, 2026, 06:38 PM UTC
This incident has been resolved.
- postmortem Oct 02, 2026, 10:05 PM UTC
# Summary On September 22, 2026, three of our outbound network gateways in one region became unhealthy and began rejecting a small percentage of new outbound connection attempts. The issue was resolved the same day problem completely with no further recurrence or follow-on impact.The underlying cause was a saturated internal traffic-monitoring buffer on those gateways, which caused their health checks to respond too slowly and get marked unhealthy. # Impact * Only three gateway instances in one region were affected; other regions and other gateways in the same region were not impacted. * There was no cascading failure , the issue was fully resolved with a single round of gateway restarts, and no follow-on issues occurred afterward. # Root Cause Our outbound network gateways run an internal component that monitors network traffic for operational visibility. Under sustained high traffic, this component's internal buffer became saturated, which created backpressure severe enough to slow down the gateway's own health-check responses beyond the timeout our load balancer allows. Once that timeout was exceeded, the load balancer marked the affected gateways unhealthy and began rejecting a portion of new connections destined for them. Actual packet forwarding continued to work normally throughout the incident. This was specifically a health-check and monitoring-buffer issue, not a failure of the gateways' core forwarding capability. # Resolution The three affected gateway instances were manually restarted, which cleared the saturated monitoring buffer and immediately restored normal health-check behavior. Connection rejections stopped as soon as the restart completed, and no further intervention was needed. # Preventive Actions | **Action** | | --- | | Replace the current traffic-monitoring approach on all gateways with a more efficient, lower-overhead implementation that cannot saturate in this way, and that alerts well before any customer-visible impact. | | Implement automatic, periodic replacement of long-running gateway instances on a fixed schedule, so no instance can accumulate the kind of age-related degradation seen in this incident. | | Migrate gateway instances to network-optimized infrastructure with substantially more dedicated network capacity, better suited to sustained high-throughput traffic. | | Enhance existing monitoring buffer's utilization while the longer-term replacement is rolled out. | | Enhance alerting specifically on gateway health-check response latency, to catch early signs of degradation before it becomes customer-visible. |