Tinybird incident

Degraded performance in the AWS us-east public cluster.

Minor Resolved View vendor source →

Tinybird experienced a minor incident on June 24, 2026 affecting Copy and Query API and 1 more component, lasting —. The incident has been resolved; the full update timeline is below.

Started
Jun 24, 2026, 03:43 PM UTC
Resolved
Jun 24, 2026, 03:43 PM UTC
Duration
Detected by Pingoru
Jun 24, 2026, 03:43 PM UTC

Affected components

CopyQuery APICopyQuery APIEvents API

Update timeline

  1. investigating Jun 22, 2026, 04:51 PM UTC

    Status: Investigating There’s a performance degradation in the public cluster on AWS us-east. Affected components Query API (Degraded performance) Copy (Degraded performance) Events API (Degraded performance)

  2. investigating Jun 22, 2026, 05:49 PM UTC

    Status: Investigating We keep working on identifying the root cause. There's impact in both reads and writes, and in all clusters in the region as the incident is affecting a shared component. We are storing the data that gets to the Events API to retry the inserts later. Affected components Copy (Partial outage) Events API (Partial outage) Query API (Partial outage)

  3. identified Jun 22, 2026, 06:52 PM UTC

    Status: Identified We're still working on stabilising the region. We're struggling with Zookeeper, in charge of coordination between ClickHouse replicas, which is why write are a lot more affected than reads. We've just tweaked a couple parameters that hope will aid the situation. Affected components Events API (Partial outage) Query API (Degraded performance) Copy (Partial outage)

  4. identified Jun 22, 2026, 09:30 PM UTC

    Status: Identified We've already identified the root cause and we're working through the fix. Real-time ingestion is now fully available, we're working on ingesting past data that hasn't been processed yet. Affected components Events API (Partial outage) Query API (Degraded performance) Copy (Partial outage)

  5. monitoring Jun 22, 2026, 09:56 PM UTC

    Status: Monitoring We've identify the root cause and we are recovering the retries from the HFI ingestion at the moment. The platform should be stable now Affected components Events API (Degraded performance) Query API (Degraded performance) Copy (Degraded performance)

  6. monitoring Jun 22, 2026, 10:05 PM UTC

    Status: Monitoring We keep ingesting the delayed real time data Affected components Events API (Operational) Query API (Operational) Copy (Degraded performance)

  7. monitoring Jun 23, 2026, 10:26 AM UTC

    Status: Monitoring We're continuing to work through a backlog of delayed ingestion on our AWS US East 1 shared cluster following yesterday's degradation. Our coordination layer was resized last night and the bulk of the impact has cleared. Recovery is taking longer than expected, so affected workspaces may continue to see ingestion lag while we work through the remaining queue. Affected components Events API (Operational) Query API (Operational) Copy (Degraded performance)

  8. monitoring Jun 23, 2026, 02:03 PM UTC

    Status: Monitoring The bulk of the earlier retries has now been recovered, and no new ingestion errors are being introduced on the customer side. Recovery is progressing more slowly than we'd like because we are currently being rate-limited by our upstream storage provider, which is limiting how quickly we can drain the remaining retries. No data has been lost, all affected inserts are being safely retried. We'll provide the next update as soon as we see meaningful progress on either front. Affected components Events API (Operational) Query API (Operational) Copy (Degraded performance)

  9. monitoring Jun 24, 2026, 07:54 AM UTC

    Status: Monitoring We are still getting rate-limited by our upstream storage provider but we are draining retries faster than before because of some optimizations we have made on the system. We are working on accelerating them as much as possible but we still have some retries in the queue. No data has been lost, all affected inserts are being safely retried. Affected components Events API (Operational) Query API (Operational) Copy (Degraded performance)

  10. resolved Jun 24, 2026, 03:43 PM UTC

    Status: Resolved All data has been recovered and the system health is back to normal Affected components Events API (Operational) Query API (Operational) Copy (Operational)