Harness experienced a major incident on May 27, 2026 affecting Continuous Integration Enterprise(CIE) - Windows Cloud Builds and Continuous Integration Enterprise(CIE) - Windows Cloud Builds and 1 more component, lasting 13h 13m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating May 27, 2026, 12:35 PM UTC
We are currently investigating this issue.
- identified May 27, 2026, 01:43 PM UTC
The issue has been identified and a fix is being implemented.
- identified May 27, 2026, 02:06 PM UTC
We are continuing to work on a fix for this issue.
- identified May 27, 2026, 03:20 PM UTC
Harness is continuing to investigate and implement changes to fully restore functionality. At this time, we are still seeing some CI Builds intermittently fail.
- identified May 27, 2026, 04:32 PM UTC
Harness has implemented a change and are seeing failures reduced, but are continuing to work on completing mitigation. Customers can expect executions to succeed more frequently. Prod1/Prod2 are back to normal.
- identified May 27, 2026, 05:54 PM UTC
Harness is currently implementing a failover to mitigate the intermittent issue for some customers
- monitoring May 27, 2026, 11:53 PM UTC
A fix has been implemented and we are monitoring the results.
- resolved May 28, 2026, 01:48 AM UTC
This incident has been resolved.
- postmortem Jun 04, 2026, 03:30 PM UTC
### **Summary** Between May 27 and June 1, 2026, some Harness CI customers experienced pipeline execution failures with the error: `failed to call LE.RetryStartStep: context deadline exceeded` ### **Impact** Affected customers saw intermittent CI pipeline failures during step execution. Existing pipeline definitions, customer data, source code, and artifacts were not impacted. ### **Root Cause** The root cause was a deadlock in the Light Engine logging path. When the log service returned an error, the Light Engine log writer attempted to reacquire a mutex it already held. This caused the Light Engine process to freeze, which led to step execution timeouts and pipeline failures. ### **Mitigation and Resolution** Harness Engineering took multiple mitigation steps during the incident, including: * Rolled back affected runner versions where needed * Increased relevant timeout configurations * Reduced log-service load and latency * Temporarily disabled the affected livelog streaming path * Migrated selected workloads across regions and infrastructure providers * Pinned a fixed Light Engine version through runner configuration The final fix addressed the mutex deadlock in the Light Engine log writer and prevented the same lock from being reacquired while already held. ### **Prevention and Follow-Up Actions** Harness is taking the following actions to reduce recurrence risk: * Improve deadlock detection in critical concurrent code paths * Strengthen error handling for log-service interactions * Add better monitoring for Light Engine process health * Improve safeguards around logging-path failures * Continue reviewing runner rollout and validation processes