Harness experienced a minor incident on June 26, 2026 affecting US - app.traceable.ai / api.traceable.ai and APAC - app.apac.traceable.ai / api.apac.traceable.ai and 1 more component, lasting 11m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating Jun 26, 2026, 07:20 AM UTC
We are currently experiencing intermittent login issues with the Traceable cluster due to an issue with our login provider. We are actively working with the provider to resolve the issue. This incident does not impact data ingestion, and all ingestion pipelines continue to function normally. We will provide updates as we have more information.
- identified Jun 26, 2026, 07:20 AM UTC
The issue has been identified and a fix is being implemented.
- monitoring Jun 26, 2026, 07:29 AM UTC
Issue has been fixed and we are closely monitoring
- monitoring Jun 26, 2026, 07:29 AM UTC
We are continuing to monitor for any further issues.
- resolved Jun 26, 2026, 07:31 AM UTC
This incident has been resolved.
- postmortem Jun 30, 2026, 05:13 PM UTC
## **Summary** _Between **06:50 UTC and 07:20 UTC on June 26, 2026**, customers experienced intermittent login failures while accessing Traceable environments across multiple US clusters. The incident was traced to an issue with the external authentication provider \(Auth0\), where elevated socket timeouts caused increased login latency and authentication request failures. Login functionality gradually recovered as the upstream issue stabilized, and the incident was resolved after successful login validation across multiple impacted clusters._ ## **Root Cause** _The root cause was an issue with the external authentication provider \(Auth0\), which experienced elevated socket timeouts while processing authentication requests. These upstream timeouts increased login latency and caused intermittent authentication failures across multiple US-region clusters. Since authentication requests depended on the external provider, affected login attempts failed despite Traceable platform services remaining healthy. The incident was resolved once the upstream authentication service recovered and login requests consistently completed successfully._ ## **Impact** _Starting at approximately 06:50 UTC, customers experienced intermittent login failures when accessing Traceable environments across multiple US clusters. The issue affected user authentication, preventing some users from accessing the platform while underlying application services remained operational. Login functionality progressively recovered during the incident, and normal authentication was restored by 07:20 UTC._ ## **Remediation** _The engineering team worked with the external authentication provider while continuously monitoring authentication health across affected clusters. Login functionality was validated through platform metrics and manual verification across representative environments. After confirming consistent authentication success across impacted clusters, the incident was declared resolved._ ## **Action Items** _To prevent such issues going forward, Harness will,_ _Increase authentication resilience: Evaluate improvements to authentication request handling, including timeout tuning, retry strategies where appropriate, and graceful degradation for transient upstream failures._