UiPath experienced a minor incident on June 18, 2026 affecting Context Grounding and Agentic Orchestration and 1 more component, lasting 6h 25m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating Jun 18, 2026, 03:00 PM UTC
We are investigating reports of degraded performance impacting workflows for Maestro in east US. Impact: Users may experience intermittent failures or stuck workflows. Our teams are working to identify the cause and will share more details as the investigation progresses.
- monitoring Jun 18, 2026, 04:39 PM UTC
We have resolved the situation that was causing the degraded performance for Maestro workflows. We had to scale up the capacity for the underlying services. We are actively monitoring the backlog of requests for latencies and will take the action as needed.
- monitoring Jun 18, 2026, 05:46 PM UTC
We have resolved the situation that was causing the degraded performance for Maestro workflows. We had to scale up the capacity for the underlying services. We are actively monitoring the backlog of requests for latencies and will take the action as needed.
- monitoring Jun 18, 2026, 05:48 PM UTC
Mitigation has been applied and performance is improving for the issue that impacted [impacted functionality / user symptom] in Maestro across US. Impact: Users may still experience degraded perfornance but this should continue to improve. We are monitoring closely to ensure stability.
- monitoring Jun 18, 2026, 07:13 PM UTC
We have resolved the situation that was causing the degraded performance for Maestro workflows. Impact: Users may still experience degraded performance, but this should continue to improve. We are monitoring closely to ensure stability.
- monitoring Jun 18, 2026, 09:03 PM UTC
We have resolved the situation that was causing the degraded performance for Maestro workflows. Impact: Users may still experience degraded performance, but this should continue to improve. We are monitoring closely to ensure stability.
- resolved Jun 18, 2026, 09:25 PM UTC
The issue has been resolved and Maestro workflow performance has returned to expected levels after degraded performance impacted Maestro workflow in US. Impact: No ongoing user impact. Agent executions that failed will have to be retried manually.
- postmortem Jul 09, 2026, 11:10 PM UTC
## Customer Impact Between 14:03 and 21:26 UTC on June 18, 2026, and between 17:02 and 23:59 UTC on June 23, 2026, a subset of customers in the U.S. Region, using Maestro and Agent executions—including Conversational Agents and some previously deployed Autonomous Agents—experienced performance degradation and workflows getting stuck in pending state. The two events, a week apart, shared the same underlying root cause. In each case, new workflow operations were stabilized before all pending workflows were fully recovered. During these windows, affected workflows could experience higher-than-normal latency, fail intermittently, or remain stuck in a pending state until recovery processing completed. ## Root Cause The incident was caused by our workflow execution backend exhausting its configured memory capacity and restarting repeatedly. Those restarts disrupted workflow state processing, causing some workflow executions to experience latency, fail intermittently, or remain pending. The primary driver of memory growth was an experimental multi-cluster \(active/passive\) replication feature for disaster recovery purpose. This feature accumulated large, long-lived workflow-history objects in memory faster than they could be released. Each time a history-processing component exceeded its memory limit and restarted, it had to reload its in-flight work—including the replication workload — which drove memory back over the limit and triggered another restart. This self-reinforcing cycle continued until responders intervened, and it was not fully resolved by the initial mitigation applied on June 18, which is why the condition recurred on June 23. ## Detection Both events were detected quickly through automated monitoring, including synthetic tests that continuously exercise the platform and alerts for pending/stuck workflow tasks. At the time of the incidents, there was no dedicated alerting on the memory saturation and restart behavior of the affected components, which lengthened the time required to identify the true root cause. Public status updates began after customer impact was confirmed. ## Response and Recovery A cross-team engineering bridge was established to investigate and mitigate the incidents. Early mitigation attempts focused on scaling backend infrastructure and restarting the affected history-processing components. These provided limited relief because they did not address the underlying memory-growth mechanism. Once the memory-retention path was pinpointed through diagnostic heap analysis, the incidents were mitigated by two actions: **increasing the memory capacity allocated to the affected history-processing components**, and **disabling the multi-clusters replication feature** that was the main contributor to the memory growth. After these changes, new workflow operations returned to normal and the region's capacity to process workflows was restored. Workflows that had become stuck in a pending state were recovered by the engineering team where possible. A small number of workflows that customers had already canceled were not replayed, and a limited set of Agent executions that had failed during the incident had to be retried manually by customers. Both incidents were marked resolved after all recoverable workflows had been processed successfully. ## Follow-Up To prevent recurrence and improve detection, we are implementing the following: 1. Right-size the memory capacity and limits for the affected components, and add automatic memory-cap safeguards so the components stay within safe bounds under load. 2. Redesign and re-validate the cross-region replication feature—including load testing with replication enabled—before it is re-enabled, so it can support disaster-recovery needs without causing memory pressure. 3. Add repeatable memory-diagnostic collection capability and operational runbooks for the service, so memory analysis and mitigation can be performed quickly during future incidents.