Harness experienced a major incident on August 6, 2026 affecting Software Engineering Insights (SEI) / AI DLC Insights and Software Engineering Insights (SEI) /AI DLC Insights and 1 more component, lasting 1h 30m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating Aug 06, 2026, 02:30 PM UTC
We are currently investigating a Harness component that is experiencing issues. We are working to identify the cause and restore normal operations as soon as possible.
- investigating Aug 06, 2026, 02:32 PM UTC
We are continuing to investigate this issue.
- identified Aug 06, 2026, 03:46 PM UTC
The issue has been identified and a fix is being implemented.
- resolved Aug 06, 2026, 04:01 PM UTC
This incident has been resolved.
- postmortem Aug 10, 2026, 06:24 PM UTC
## Summary Customers on Prod1, Prod2, and Prod3 \(US\) clusters experienced failures when loading SEI 2.0 dashboards on August 6, 2026, from 7:22 AM PDT to 9:03 AM PDT. Customers calling the SEI 2.0 API also experienced similar failures. No customer data was lost, and ingestion of all integration data continued to work uninterrupted. SEI customers using 1.0 were not impacted. ## Root Cause The incident was caused by resource exhaustion on the nodes serving queries. This resource degradation developed in a pattern that did not cross our existing alerting thresholds early enough to provide sufficient warning or allow mitigation before customer impact occurred. ## Impact Customers on Prod1, Prod2, and Prod3 \(US\) clusters were unable to load SEI 2.0 dashboards during the incident window. **Duration:** August 6, 2026, 7:22 AM PDT – 9:03 AM PDT \(~1 hour 41 minutes\) ### What was not impacted? * Data ingestion and processing * SEI 1.0 customers * Integrations and metadata flows No customer data was lost. ## Remediation Upon identifying the root cause, our team took immediate corrective action by adding capacity to restore the affected systems. Services were fully recovered, and all dashboards resumed normal operation at 9:03 AM PDT. ## Action Items To prevent from such issues happening again, Harness is/has Proactively added capacity updates have been applied to prevent this issue from recurring #### 2. Enhanced Monitoring and Alerting Additional monitoring and alerting have been put in place to detect anomalies early, focused on a leading indicator, which in this case was thread pool exhaustion, before they can impact dashboard availability and data rendering. #### 3. System Patch in Progress We are working with our vendor to apply a patch to remediate this and similar issues completely.