Keeper incident

EU KeeperPAM router and gateway connections

Notice Resolved View vendor source →

Keeper experienced a notice incident on July 31, 2026, lasting —. The incident has been resolved; the full update timeline is below.

Started
Jul 31, 2026, 08:00 PM UTC
Resolved
Jul 31, 2026, 06:00 AM UTC
Duration
Detected by Pingoru
Jul 31, 2026, 08:00 PM UTC

Update timeline

  1. resolved Jul 31, 2026, 08:00 PM UTC

    Connections established through the Keeper Router and Gateway in the EU region is failing. Additional details will be posted in the postmortem report.

  2. postmortem Jul 31, 2026, 08:00 PM UTC

    # Post-Incident Report **July 31, 2026** ## Summary On July 30, 2026 at 11:33 PM CT, the KeeperPAM connection routing service in the EU region started throwing errors. The Router/Gateway errors lasted approximately 12 hours and 45 minutes, with full service restored at approximately 12:18 PM CT on July 31. During this window, the main Keeper EU platform — including vault access, authentication, and all other Keeper services — remained fully operational throughout. ## What Happened On July 22, a routine deployment to our connection routing service \(“Keeper Router”\) included an updated dependency that changed how the service retrieves its startup configuration. The endpoint is related to FIPS-specific endpoints. A misconfiguration in the endpoint caused EU region to look for a FIPS endpoints, which are not available. As a result, when service containers in the EU region restarted following the deployment, they were unable to retrieve their startup configuration and entered a failed state. Two factors allowed this failure to go undetected for over a week. First, existing containers continued serving traffic while replacement containers silently failed to start, so there was no immediate customer impact from the July 22 deployment. Second, the configuration-load failure was logged at INFO level rather than ERROR or CRITICAL, so no alerts were generated. On the night of July 30, the last healthy containers cycled out, the service's load balancer had zero healthy targets, and the endpoint began returning HTTP 503 errors. Automated health-check monitoring detected the outage within seconds and paged our on-call team. ## What We're Changing 1. **Alerting on failed container rollovers.** We are deploying CloudWatch alarms across all regions and environments that fire when container tasks repeatedly fail to start, to detect silent deployment failures. 2. **Log severity for startup failures.** Any failure to load required startup configuration will now log at ERROR/CRITICAL severity instead of INFO which trigger the proper actionable alerts. We apologize to our EU customers for the disruption.