Harness incident
Dashboard aggregator experiencing timeouts across prod2 since 15:04 UTC (Sept 21)
Harness experienced a minor incident on September 21, 2026 affecting Custom Dashboards, lasting 1h 53m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating Sep 21, 2026, 04:39 PM UTC
We are currently investigating this issue.
- identified Sep 21, 2026, 05:12 PM UTC
The issue has been identified and a fix is being implemented.
- monitoring Sep 21, 2026, 06:27 PM UTC
A fix has been implemented and we are monitoring the results.
- resolved Sep 21, 2026, 06:33 PM UTC
This incident has been resolved.
- postmortem Sep 29, 2026, 10:39 PM UTC
## Summary On Sep 21, 2026, customers in Prod 2 experienced degraded and then largely unavailable deployment dashboards for approximately three hours. The issue was caused by a set of analytics queries behind those dashboards reading far more data than they needed, which exhausted the capacity of the database serving them. A rolling restart of the affected service restored dashboards fully, and the platform has been stable since. ## Customer Impact Dashboard pages loaded slowly or timed out between 14:48 and 18:11 UTC. Pipeline execution and deployment recording run on separate infrastructure and continued to operate normally. No data loss or corruption occurred. ## Root Cause The dashboard queries were missing a time-range filter on the column used to split the largest table into weekly segments, so each query read all historical data instead of only the requested period. Correctness was never affected — the dashboards returned the right results; only the amount of data read to produce them was excessive. As data volume grew, the cost of those reads exceeded what the database could serve, and retries of the failing queries added further load, so the condition did not clear until the service was restarted. ## Mitigation * Performed a rolling restart of the affected service, which stopped the retry cycle and allowed the database to complete its outstanding work * Confirmed dashboard latency and database connection usage returned to normal levels ## Next Steps To prevent recurrence, Harness will: * Add the missing time-range filter to the affected dashboard queries so each reads only the requested period * Replace immediate retries with bounded, backing-off retries * Improve detection by adding alerting on database connection usage and per-query execution cost