Thycotic incident

Platform: Availability Failures and Degraded Performance

Major Resolved View vendor source →

Thycotic experienced a major incident on September 16, 2026 affecting Platform, lasting 1h 55m. The incident has been resolved; the full update timeline is below.

Started
Sep 16, 2026, 04:37 PM UTC
Resolved
Sep 16, 2026, 06:32 PM UTC
Duration
1h 55m
Detected by Pingoru
Sep 16, 2026, 04:37 PM UTC

Affected components

Platform

Update timeline

  1. investigating Sep 16, 2026, 04:37 PM UTC

    We are investigating an issue affecting login and access to the Delinea Platform in the US region. Some users may be unable to log in or may experience slow performance when accessing their tenants. Our engineering team is actively engaged. We will provide an update as soon as more information is available.

  2. identified Sep 16, 2026, 04:50 PM UTC

    The issue has been identified and a fix is being implemented.

  3. monitoring Sep 16, 2026, 05:02 PM UTC

    A fix has been implemented and we are monitoring the results.

  4. monitoring Sep 16, 2026, 05:22 PM UTC

    We are continuing to monitor for any further issues.

  5. resolved Sep 16, 2026, 06:32 PM UTC

    This incident has been resolved.

  6. postmortem Sep 24, 2026, 09:50 PM UTC

    ## Incident Overview On September 16, 2026, Delinea Platform customers in the US region experienced login failures and severe slowness when signing in or accessing their tenants. Affected users may have seen "upstream request timeout" errors on the sign-in screen or experienced long delays when loading their tenant. * **Start:** September 16, 2026 approximately 11:40 AM EDT \(UTC-4\) * **End:** September 16, 2026 approximately 1:02 PM EDT \(UTC-4\) ## Root Cause and Remediation The Delinea Platform relies on a caching service to support login, authentication, and other identity-related operations. Over the preceding weeks, demand on this service in the US region had grown steadily with increasing platform usage, leaving limited operating headroom. In addition, authentication services were holding connections to the caching service open over time rather than releasing them, which gradually consumed its available capacity. A configuration that distributes read traffic across an additional caching resource, which would have provided extra headroom, was not active at the time. On September 16, normal traffic volumes pushed the caching service beyond its operating capacity, causing slow and failed requests for login and identity operations across the US region. This incident involved the same caching infrastructure as the [August 28, 2026 incident](https://status.delinea.com/incidents/1s7gw9kyjwyq). That incident was triggered by an unusually large volume of synchronization activity and was addressed by increasing the caching service's memory capacity. This incident occurred under normal traffic conditions, which showed that compute capacity and connection handling also needed to be addressed. Our engineering team restarted the affected authentication services, which released the accumulated connections and restored login functionality. The caching service was also scaled to a higher tier that doubles its compute capacity, and we continued to monitor the service until full recovery was confirmed. Since the incident, caching service utilization has remained within normal operating thresholds, and read traffic is now distributed across an additional caching resource. ## Preventative Actions Following the August 28, 2026 incident, we committed to several improvements to this caching infrastructure. Below is the progress on each, along with additional actions identified from this incident. Work on all items is ongoing. **Progress on actions from the August 28 incident** * **Increase the capacity and resiliency of our caching infrastructure.** The US caching service has been scaled to a higher tier that doubles its compute capacity. Read traffic is now distributed across an additional caching resource. We are continuing this work by moving token-related operations to dedicated caching infrastructure for further headroom. * **Correct monitoring thresholds to detect capacity pressure proactively.** During this incident, our monitoring did identify capacity pressure ahead of customer impact, but the alerts were not actioned in time. We are continuing this work by removing non-actionable alerts for core authentication services and requiring immediate engineering engagement on the first alert. * **Limit connection-recovery retries.** This work has been expanded to cover how authentication services manage caching connections overall, including connection release and timeout settings, so capacity is not gradually consumed over time. * **Apply rate limiting to large synchronization jobs.** Rate limiting is being implemented. We are also reducing unnecessary high-volume requests from integrated services to the caching layer. **Additional actions from this incident** * Implement scheduled, automated rolling restarts of authentication services as an interim safeguard until the connection handling improvements are deployed. We recognize this is not the first sign-in disruption on the Delinea Platform in recent weeks, and we are not treating it lightly. Beyond the actions above, Delinea leadership has made our Platform reliability an executive-sponsored priority, ahead of feature work where the two compete. That includes better detection of issues affecting critical workflows such as sign-in, and planning capacity ahead of growth. We sincerely apologize for the disruption this caused to your operations and will continue to share our progress.