Okta experienced a major incident on June 22, 2026 affecting Core Platform, lasting 16d 17h. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating Jun 22, 2026, 09:22 PM UTC
At 6/22/2026 2:22 PM PT, the team became aware of an issue with our Core Platform service affecting customers in EU Cell 1. During this time, users may experience intermittent failures and high latency when logging in or accessing the service. Our team is actively investigating this issue and is working to mitigate it. We will provide another update within the next 15 minutes, or sooner if additional information becomes available. Affected cells: okta.com:1
- monitoring Jun 22, 2026, 10:02 PM UTC
The Okta Engineering team has observed that the errors have been reduced. We are monitoring the situation closely and will provide further updates as we move toward full mitigation.
- resolved Jun 22, 2026, 10:33 PM UTC
Internal monitoring confirmed that the EU1 environment has fully recovered and is operating normally. A full Root Cause Analysis (RCA) will be made available within 5 days. We apologize for any disruption this may have caused.
- resolved Jun 30, 2026, 02:43 PM UTC
We sincerely apologize for any impact this incident has caused to you, your business, and your customers. At Okta, trust and transparency are our top priorities. Outlined below are the facts regarding this incident. We are committed to implementing improvements to the service to prevent future occurrences of this incident. Detection and Impact: On June 22nd, 2026, at 2:08 PM PDT, Okta internal monitoring alerted our team to errors in the EU environment. During this period, administrators and end users may have experienced intermittent 500 Internal Server Error and 429 Rate Limit responses when attempting to access Okta services. Root Cause Summary: The incident was triggered by an unexpected scaling down of the application cluster during scheduled maintenance of the environment. This change accidentally overwrote the custom capacity settings for several of our applications. This caused the system to automatically revert to its lowest default capacity, drastically reducing the amount of user traffic it could handle simultaneously. As a result, the remaining servers were immediately overwhelmed by the normal volume of user requests, leading to system slowdowns, errors, and "too many requests" blockages for our customers in the impacted cell. Remediation Steps: Okta's engineering team identified the root cause issue at 2:44 PM PDT. The team mitigated the issue by scaling up resource capacity, completely stabilizing service health back to normal levels by 2:48 PM PDT. Preventative Actions: In order to prevent similar incidents from happening again, Okta is currently reviewing the following: - Strengthen guardrails by improving our validation tools and multi-environment review processes to automatically detect and flag unexpected server capacity changes before they’re applied. - Accelerate system monitoring by implementing enhanced monitoring thresholds to instantly detect and report infrastructure capacity drops before they can impact end users. - Optimize failover protocols by auditing critical alert paths for maximum reliability while establishing a clear, structured communication plan for when automated backup systems are triggered. These learnings work to prevent similar incidents from happening again. Duration (# of minutes):40 minutes