Is Ambra down?
Last checked 2m agoNo incidents right now.
Ambra is operational right now. Last checked 2m ago; the most recent incident resolved 59d ago.
Real-time Ambra status, recent outages, and incident history — pulled directly from Ambra's official status page at https://ambra.statuspage.io every 5 minutes. Pingoru tracks 3 Ambra services and has captured 1 incident in the last 90 days (98.89% uptime). Get email, Slack, Discord, or webhook alerts the moment Ambra reports a new incident — free for 5 monitors, no credit card.
Recent outages & incidents
Past 90 days- Web ServicesImage ProcessingImage Viewing
Timeline · 6 updates
- investigating · Jul 20, 2026, 01:11 PM UTC
We have received reports of log in issues on the InteleShare platform. Users appear to be unable to log in, or initial log in is slow. Engineering teams are currently investigating. Additional information will be provided as soon as it is available.
- investigating · Jul 20, 2026, 01:41 PM UTC
Our Engineering teams are actively investigating and working to identify the root cause. We understand the urgency and we appreciate your patience as we work to address the issue.
- investigating · Jul 20, 2026, 02:04 PM UTC
Our engineering teams have not yet identified the root cause. We are continuing to investigate and will provide further updates as soon as possible.
- monitoring · Jul 20, 2026, 02:29 PM UTC
At this time, our engineering team has made a configuration change to our system backend to improve stability. Since implementing this change, the website’s frontend has been operating reliably. The root cause has been identified, and we will continue to monitor system performance.
- resolved · Jul 20, 2026, 03:00 PM UTC
The incident has been fully resolved and service is back to normal levels. Our team will be conducting a root cause analysis and sharing as soon as possible. We will continue to monitor the situation to ensure there are no further issues.
- postmortem · Jul 21, 2026, 08:44 PM UTC
**What Changed** As part of the scheduled weekend maintenance window, the InteleShare infrastructure was switched to a more modern autoscaling and container-scheduling approach using Karpenter. Karpenter selects the most cost-efficient node/instance type for the current cluster load, replacing a prior approach that effectively over-provisioned compute without the team realizing it. The intent was more efficient and cost-effective use of cloud resources, faster scaling in response to traffic, and access to a wider variety of instance types to avoid AWS capacity constraints. **What Failed** InteleShare workload has relatively low CPU and memory requirements but handles a very high volume of concurrent network connections. Once the workload moved to Karpenter's "right-sized" scheduling, several instances of that workload's container were placed on smaller nodes than before \(moved from a 2XL instance down to an XL, roughly half the CPU\). The Linux kernel's connection-tracking table sizes its maximum entries based on the node's CPU and memory. Because the smaller nodes had a much lower ceiling, and because multiple instances of the high-connection workload were scheduled onto the same nodes, the combined connection count exceeded the conntrack limit once Monday-morning production load hit. Once the limit was reached, new connections on the affected nodes experienced delays or timeouts - and because the limit is enforced at the OS level, it affected other containers co-located on the same nodes as well. **Why It Wasn't Caught Earlier** The change had already been running in UAT and all other pre-production environments, which the team could observe before rollout, and it behaved correctly there. The issue only manifested under the connection volume and concurrency pattern of full U.S. production traffic - UAT volume is orders of magnitude lower and does not reproduce the same ordering/concurrency of requests. Simulating true production-scale concurrent-connection load in a lower environment is difficult and, at current test capacity, was not something the team could reasonably reproduce ahead of time. **Impact** Clients experienced degraded performance - delayed or timed-out connections - for approximately two to two and a half hours on Monday morning \(roughly 9:37 a.m. Eastern into the 10-11 a.m. hour\), concentrated around peak login/usage volume. **Resolution** The on-call platform engineer was paged automatically when the platform-wide Code Blue was triggered. The team identified the affected application tier and reverted that specific component back to the prior scheduling configuration; the rest of the Karpenter rollout \(roughly 85%\) remains in place. The revert was performed before the underlying root cause \(the conntrack ceiling\) was fully understood - root cause was confirmed afterward through investigation.
Latest: **What Changed** As part of the scheduled weekend maintenance window, the InteleShare infrastructure was switched to a more modern autoscaling and container-scheduling approach usi…
-
- InteleShare Incident ResolvedStarted Jul 20, 2026, 01:11 PM UTC · Resolved Jul 20, 2026, 03:00 PM UTC · 1h 49m