Harness incident

Prod2 was intermittently unavailable

Notice Resolved View vendor source →

Harness experienced a notice incident on September 5, 2026 affecting Platform, lasting 46s. The incident has been resolved; the full update timeline is below.

Started
Sep 05, 2026, 09:40 AM UTC
Resolved
Sep 05, 2026, 09:40 AM UTC
Duration
46s
Detected by Pingoru
Sep 05, 2026, 09:40 AM UTC

Affected components

Platform

Update timeline

  1. investigating Sep 05, 2026, 09:40 AM UTC

    We are currently investigating this issue.

  2. resolved Sep 05, 2026, 09:40 AM UTC

    This incident has been resolved.

  3. postmortem Sep 09, 2026, 06:22 PM UTC

    ## **Summary** Between 12:34am PST and 12:38am PST on 5th September, the Delegate service manager experienced some elevated exceptions when attempting to write to the database. Consequently, delegate connections were dropped, causing them to disconnect. Delegate automatically re-attempts registration back to the `delegate service manager` and majority of the delegates got connected back after the incident. For Docker and ECS delegates the automatic restart is not enabled unless these delegates have health monitoring enabled. For these delegates a manual restart is needed and was recommended. Post restart the delegate would re-connect and the issue was resolved. ## **Root cause** On Prod2 cluster we identified a performance bottleneck in the delegate service that, under certain conditions, can increase database write latency and delay heartbeat processing which leads to delegates being disconnected. ## **Impact** All K8s delegates and \`Docker/ECS\` delegates got connected back immediately within 4 mins and started to function normally. The impact can be scoped to those specific types of delegates that didn’t have health monitoring enabled. ## **Remediation** * Immediate: We have added additional monitoring and increased resources for handling the influx of traffic. * Permanent: We have identified a hotspot in the code that can cause high latency when writing to a database which we are actively working on resolving. ## **Action Items** To prevent such issues from happening again, Harness will work on the following: 1. Increased targeted monitoring and alerting to initiate timely mitigation and prevent this from happening again. 2. Fix the identified delegate service managers database client reconnect failures 3. Fix the hotpots that can cause query latency.