Keeper Outage History
Keeper is up right nowKeeper had 21 outages in the last 2 years totaling 29h 45m of downtime — averaging 0.9 incidents per month.
There were 21 Keeper outages since September 24, 2024 totaling 29h 45m of downtime. Each is summarised below — incident details, duration, and resolution information.
Resolved: Keeper Gateway and KSM API errors
Timeline · 2 updates
- identified Jul 08, 2026, 06:51 PM UTC
We are aware of Keeper Gateway and KSM connection errors due to "invalid client version" responses. We have identified the issue and it will be resolved in a few minutes.
- resolved Jul 08, 2026, 07:14 PM UTC
API errors have been resolved. The Keeper Gateway and KSM client version is no longer returning an error.
Direct record sharing enforcement issue
Timeline · 2 updates
- identified Jun 19, 2026, 03:05 PM UTC
We are aware of an issue with direct record sharing for enterprise users as a result of yesterday's Backend API 18.1.3 update. We will issue a patch this morning.
- resolved Jun 19, 2026, 11:00 PM UTC
The issue was resolved at 12:30PM PST.
API errors during maintenance operation
Timeline · 1 update
- resolved Apr 01, 2026, 01:40 AM UTC
On March 30 at 8:00 PM PST, Keeper performed a scheduled minor upgrade to our multi-region AWS RDS clusters as part of routine maintenance. The upgrade itself completed in approximately 3 minutes, followed by a standard application restart process that took about 15 minutes. While this process is typically non-disruptive, recent changes to service health checks caused unexpected customer impact during the restart window. This maintenance was not communicated in advance due to an internal miscommunication and the expectation of no impact. We apologize for the disruption and are taking steps to ensure all future maintenance—regardless of expected impact—is communicated ahead of time, as well as reviewing our health check configurations to prevent similar issues.
Resolved: KeeperPAM Connection Errors in US Data Center
Timeline · 5 updates
- investigating Mar 23, 2026, 05:07 PM UTC
We are investigating errors in establishing KeeperPAM connections in the US Data Center.
- identified Mar 23, 2026, 06:47 PM UTC
We have identified the cause of the KeeperPAM connection errors which are related to a networking issue within the AWS ECS environment. We are actively troubleshooting with the AWS support team and will update the ticket as soon as there is a status update.
- monitoring Mar 23, 2026, 08:03 PM UTC
A fix has been implemented and we are monitoring the KeeperPAM connection stability. We will update this case with additional details soon.
- resolved Mar 23, 2026, 09:45 PM UTC
The issue has been resolved. See postmortem page with additional details.
- postmortem Mar 23, 2026, 11:17 PM UTC
At 9:05 AM PST, alerts were triggered for issues affecting KeeperPAM connections managed through Keeper’s ECS deployments in the US-EAST region. There were no recent changes to the application or environment. Investigation identified a low-level concurrency bug in the Keeper EPM service that caused request failures under high simultaneous load. These failures led to instability in the ECS services supporting KeeperPAM connections. As a temporary mitigation, we blocked the error condition, restoring KeeperPAM connectivity by approximately 1:00 PM PST. The engineering team then developed and deployed an updated Keeper Router version to address the underlying issue and prevent EPM agents from triggering server errors. The fix was fully validated by 3:00 PM PST, at which point all KeeperPAM services were stable and operating normally.
US GovCloud Region - Login Errors Resolved
Timeline · 2 updates
- resolved Feb 19, 2026, 05:30 AM UTC
As part of our planned GovCloud region maintenance involving database upgrades, restarts of the Redshift cluster were required. As a result, downstream services required configuration updates and all services are now restored.
- postmortem Feb 19, 2026, 05:30 AM UTC
Maintenance in the GovCloud region was scheduled to begin at 8:00 PM PST with a planned duration of 30 minutes. During the maintenance window, certain services required additional reconfiguration beyond the original scope. Core services were restored within the planned 30-minute window. However, due to the extended reconfiguration activities, some users experienced API errors during login for approximately 15 minutes following service restoration.
EU data center login errors - Resolved
Timeline · 3 updates
- identified Dec 02, 2025, 10:48 PM UTC
During scheduled system updates, we have encountered unexpected NGINX 503 / 502 errors in the EU region. As of 4:50 PM CST, work is still in progress. During this time, users in the EU region will experience brief or intermittent access interruptions when using Keeper.
- resolved Dec 03, 2025, 12:47 AM UTC
The issue has been resolved. We will update the postmortem report shortly.
- postmortem Dec 04, 2025, 10:29 PM UTC
On December 2, our **EU region** experienced an unexpected outage during planned infrastructure updates. The issue began at **1:24 PM PT** and service was fully restored by **4:10 PM PT**. **What Happened** During an update to our server configuration, a region-specific configuration error caused the EU environment to begin throwing API errors. Although the underlying issue was identified quickly, each corrective change required a full autoscaling group rollout \(up to 30 minutes per cycle\) which extended the total recovery time. **Resolution** Our engineering team implemented the necessary fixes and restored service. We have also updated our infrastructure-as-code to ensure this issue cannot recur in future deployments. We apologize for the disruption and appreciate your patience while we resolved the incident.
AWS outage affecting KeeperPAM connections and scheduled rotations
Timeline · 2 updates
- identified Oct 20, 2025, 05:25 PM UTC
There are intermittent connectivity issues due to the AWS outage in US-EAST-1 affecting KeeperPAM cloud connections, scheduled rotations and discovery operations.
- resolved Oct 20, 2025, 08:31 PM UTC
Clearing this alert as most AWS services are restored.
Email delivery in US region due to AWS US-EAST outage
Timeline · 3 updates
- investigating Oct 20, 2025, 09:20 AM UTC
AWS is reporting errors in the US-EAST region that are affecting multiple services, affecting email and SMS delivery which impacts Keeper device approvals. The AWS Service Health dashboard can be monitored here: https://health.aws.amazon.com/health/status
- identified Oct 20, 2025, 10:00 AM UTC
SMS delivery services have been restored, but AWS Lambda is still down which affects our email delivery of device approvals. We recommend using the "Keeper Push" or TOTP method of device approval.
- resolved Oct 20, 2025, 03:17 PM UTC
All Keeper email delivery services in US-EAST-1 were restored around 4AM PST.
EU Region login API errors
Timeline · 3 updates
Resolved: Vault login API errors
Timeline · 1 update
- resolved Aug 21, 2025, 05:36 PM UTC
Resolved - At 9:49 AM PST, we experienced login API request failures due to errors on an RDS database in the US region. The database failover completed successfully, and all systems were fully operational by 10:00 AM PST.
SSO login errors with strict SP cert matching
Timeline · 2 updates
Vault login errors - resolved
Timeline · 4 updates
- investigating Jun 24, 2025, 03:27 PM UTC
We are investigating vault login errors
- investigating Jun 24, 2025, 03:28 PM UTC
We are continuing to investigate this issue.
- resolved Jun 24, 2025, 04:01 PM UTC
The issue has been resolved. The outage was caused by a high volume of database locks related to a large BreachWatch data update. The engineering team is tracking down the root cause and will be addressing a preventative solution today.
- postmortem Jun 24, 2025, 08:42 PM UTC
At 8:15 AM PST, our DevOps team was alerted to elevated API error rates in production. Upon investigation, we identified the following: * The production database was experiencing high table lock contention. * This was triggered by several large-scale BreachWatch data updates. * A recent backend migration had unintentionally doubled the number of parallel update operations, which exacerbated the issue. * A database failover was performed to restore service stability. The issue was fully resolved by 8:40 AM PST. The engineering team has addressed the underlying problem in the update processes and is implementing safeguards to prevent recurrence, including query optimization and improved job orchestration. We apologize for the disruption.
Login errors in the EU data center
Timeline · 4 updates
- investigating Jun 11, 2025, 12:04 PM UTC
We are investigating login errors in the EU data center.
- identified Jun 11, 2025, 12:16 PM UTC
We have identified the issue causing login errors in the EU data center. Our team is working on a resolution.
- resolved Jun 11, 2025, 12:46 PM UTC
Login errors in the EU region have been fully resolved. We will update this incident postmortem with additional information shortly.
- postmortem Jun 12, 2025, 05:09 AM UTC
At 4:37AM PST, Keeper experienced API errors in the EU region. Service was restored by 5:35AM PST following a database failover. The issue was caused by a bug in the new Device Management API, released the day before. During device registration, the system attempted to delete old devices, which created high database load for accounts with large device counts. The bug was identified, fixed, and a patch was deployed. We apologize for the disruption. To ensure uninterrupted access, we recommend enabling “Work Offline” mode for your account.
File attachment download errors for new uploaded files
Timeline · 3 updates
Vault login errors
Timeline · 2 updates
Vault Login errors - resolved
Timeline · 3 updates
- investigating Mar 06, 2025, 06:23 PM UTC
We are investigating login errors and our team is currently working on the issue.
- investigating Mar 06, 2025, 06:24 PM UTC
We are continuing to investigate this issue.
- resolved Mar 06, 2025, 06:25 PM UTC
The issue has been resolved
Keeper vault login errors - resolved
Timeline · 3 updates
- identified Feb 24, 2025, 04:48 AM UTC
We are still working on the issue which is causing login errors in the US region. Please use Keeper's "work offline" while we are working to resolve the issue.
- resolved Feb 24, 2025, 05:05 AM UTC
The issue has been resolved and root cause determined. We will update the postmortem on this issue shortly.
- postmortem Feb 24, 2025, 07:07 AM UTC
We identified a backend API and resulting database query which triggered a high database load and affected production services. We are publishing a permanent fix on Monday, Feb 24.
Login errors in US region
Timeline · 2 updates
- investigating Feb 24, 2025, 03:54 AM UTC
We are currently investigating login errors in the US region.
- resolved Feb 24, 2025, 04:28 AM UTC
The issue has been resolved.
T-mobile SMS delivery issues in US
Timeline · 2 updates
- investigating Jan 13, 2025, 10:15 PM UTC
We are currently experiencing delivery issues with T-Mobile SMS 2FA messages. While AWS works to resolve the issue, SMS-based 2FA codes will be temporarily re-routed to email for affected users.
- resolved Jan 13, 2025, 10:50 PM UTC
The SMS delivery issue with T-mobile has been resolved.