[Resolved] Test result uploads failed
Timeline · 1 update
- investigating Oct 04, 2026, 07:00 PM UTC
The permanent fix is now live in production. Test result uploads have been working since 18:03 UTC, and the underlying cause is fixed.
Tuist had 20 outages in the last 2 years totaling 4180h 44m of downtime — averaging 0.8 incidents per month.
There were 20 Tuist outages since October 31, 2025 totaling 4180h 44m of downtime. Each is summarised below — incident details, duration, and resolution information.
The permanent fix is now live in production. Test result uploads have been working since 18:03 UTC, and the underlying cause is fixed.
We have deployed a permanent fix, and the API, remote cache and login are working normally. The outage was caused by a deployment that changed how API requests are authenticated, which overloaded our servers, and that change has been reverted. If the CLI still reports authentication errors, run `tuist auth login` again. We apologize for the disruption.
This was a traffic instability of the provider's Warsaw region. We're working on with the provider to figure out what went wrong and how we can ensure this doesn't happen in the future.
Our bot protection prohibited real users to sign up. We'll add a postmortem tomorrow to prevent this from happening in the future.
The processor got disconnected and we were not alerted immediately after that happened, which lead to tests being stuck in the processing queue. We've identified the root cause and ensuring this doesn't happen again in the future.
After monitoring the registry, we observed no further issues. It is important to note that, because some packages were rebuilt, users will need to clear their cached fingerprints, as those packages now have new checksums. We will publish a detailed RCA next week with more information about the incident and its root cause.
A volume reaching its maximum capacity prevented cache resolution, triggering a cascade of failures that left the region largely non-operational. Once disk space was restored, cache resolution resumed and the region became operational again.
A failed deployment left legacy registry synchronizers writing alongside the new registry, while regional object-storage replicas sometimes returned older archive bytes. This caused missing versions, checksum mismatches, and signature failures. We paused legacy registry syncing without affecting cache availability, repaired and audited the package metadata and archives, and confirmed zero remaining metadata checksum mismatches. Production is protected by the live fix, and pull request #12120 makes authoritative archive reads durable across future deployments.
The package syncing had a bug that prevented some versions from persisting. We've fixed the bug and deployed it, forcing the resync of the packages reported by users.
The incident has been resolved, and we are working on better observability to detect incidents of this nature earlier, as well as some changes in our infrastructure so we can recover fast and reliably from these kinds of issues. We'll share a post morten shortly
Turns out the processing demand exceeded the capacity. Test results are being processed, but it just takes longer. We are looking into ways to optimize processing time.
A new version of the server has been deployed to ensure users are not blocked from accessing the cache.
Production tuist.dev returned 503 because all main server pods were crash-looping during startup. The immediate failure was a timeout in license validation against Keygen, but Keygen itself was healthy. The real issue was the production stable egress gateway. Server pod traffic is routed through a Cilium egress gateway using the fixed IP 116.202.0.10. After node churn, the gateway node label and Hetzner floating IP were not attached to any active node, so selected server traffic could not reach the public internet. We restored service by assigning the floating IP to a live general worker, labeling that node as the stable egress gateway, forcing Cilium to refresh the policy, and restarting the server deployment. tuist.dev is now healthy again. To prevent this from recurring, we are working on making the stable egress setup declarative/self-healing instead of relying on a manual node handoff, and adding monitoring for the gateway readiness and server egress path.
Some of our caching nodes exhibited intermittent failures under high load, which we've mitigated by adding additional regional nodes. We are actively working on developing and testing a new solution that we plan to deploy in a per-tenant model fashion that will self-regulate under high-load scenarios.
All stuck test results have been processed now. We're putting up an alert to catch a condition like this sooner.
Status: Investigating The test results are stuck in "in processing". We're working on resolving the issue. The rest of the features, like build insights, are not impacted. Affected components Tuist (Operational)
Status: Resolved The test results are now being processed again. We're putting the following mitigations in place, so this doesn't happen again and the issue is resolved faster: - We will be adding a redundant Mac machine for processing the xcresults remotely - We're adding additional monitoring to ensure each new release is healthy. Affected components Tuist (Operational)
Status: Investigating Cache and registry are currently unavailable. The team is investigating. Affected components Tuist (Partial outage)
Status: Resolved The incident has been resolved. Affected components Tuist (Operational)
Status: Resolved On April 7, 2026, Tuist experienced an incident that made cache endpoints unavailable globally for approximately 12 minutes, from 16:30 UTC to 16:42 UTC. The full post-mortem is available at https://community.tuist.dev/t/post-mortem-cache-and-registry-outage-april-7-2026/966 . Affected components Tuist (Operational)