Unico experienced a major incident on July 21, 2026 affecting API, lasting 1h 23m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- identified Jul 21, 2026, 04:39 PM UTC
Dear customer, We are investigating an instability in the IDCloud functionalities that affects the authentication flows in the process journey, specifically impacting the Web Journey (formerly ByUnico). Our engineering team is mobilized to resolve the issue as soon as possible. We will provide further updates shortly.
- monitoring Jul 21, 2026, 04:49 PM UTC
Dear Customer, Our engineering team has identified and implemented the necessary actions to resolve the instability in the affected IDCloud flow capabilities. At this time, the technical team is conducting a detailed analysis of our data layer and continues to monitor the environment. We reiterate that further updates will be sent shortly.
- resolved Jul 21, 2026, 06:02 PM UTC
Executive Summary and Impact Starting at approximately 12:44 PM (Brasília time) on July 21, 2026, a portion of users in the reduced onboarding journey experienced slower loading of the initial screens in the flow. The impact was observed across all environments where this journey is used, affecting perceived speed and, to a lesser extent, users' ability to complete the process during this period. Root Cause and Resolution The investigation is ongoing to confirm the definitive root cause of the incident. There was no impact on servers, network, or infrastructure — the issue was confirmed to be isolated to the on-screen experience on the user's device. As a containment measure, the responsible team applied a preventive adjustment at 1:39 PM (BRT), and loading times returned to normal within a few minutes. The analysis to validate and confirm the root cause is still being conducted by the technical team. Commitment and Next Steps As part of the continuous improvement of incident response processes, the team commits to enhancing the automatic escalation flow for responsible teams, making activation even more direct and effective from the moment an incident is opened. Additionally, loading-speed monitoring will be strengthened during the rollout of new changes to this journey, with the goal of identifying and correcting potential regressions even faster, minimizing impact to users in similar situations going forward.
- postmortem Aug 03, 2026, 02:01 PM UTC
**Summary** On July 21, 2026, between 12:44 PM and 1:43 PM \(Brasília time\), page load performance indicators for the reduced journey degraded severely, with average load time jumping from ~4.8s to ~8.9s and the p95 reaching 31.5 seconds. The incident was caused by rate limiting applied by the cloud infrastructure on connections originating from the provider, which prevented frontend assets from loading. **Impact** Users from 32 clients experienced extremely slow page loads or complete loading failures, with approximately 94 journey errors recorded. "Initialization Timeout" errors — triggered when the application fails to mount within 30 seconds — reached about 23 thousand occurrences during the period. **Root Cause** The root cause was identified in the infrastructure's network layer: the CDN provider consolidates traffic through a limited set of outbound IP addresses. A spike in new TCP connections and TLS handshakes originating from these IPs triggered the cloud infrastructure's automatic network-layer defense mechanisms, which began rate-limiting connections even before HTTP processing. Because these failures occurred at the network layer — before HTTP processing — they did not appear in standard request logs, making the problem invisible to conventional monitoring and significantly complicating diagnosis. **Resolution** The impact ceased with the natural recovery of the rate limiting applied by the cloud infrastructure. The deactivation of a component, initially flagged as a possible cause, coincided temporally with the recovery but was later confirmed to be unrelated to the incident. Configuration fixes — enabling HTTP Keep-Alive with an extended timeout and adding an allow rule for CDN IPs at the edge protection layer — were identified as follow-up items by the support teams of the providers involved. **Lessons Learned** The incident highlighted that failures at the network layer between the CDN and cloud infrastructure are invisible to standard application monitoring, creating a significant blind spot. As follow-up items, the team identified the need to implement the configuration fixes recommended by the providers and to expand observability at the asset delivery layer to proactively detect this type of degradation.