Nango incident

Records database is experie...

Minor Resolved View vendor source →

Nango experienced a minor incident on September 13, 2026 affecting Nango Cloud Health and Nango Cloud Health, lasting 9d 20h. The incident has been resolved; the full update timeline is below.

Started
Sep 13, 2026, 11:00 PM UTC
Resolved
Sep 23, 2026, 07:48 PM UTC
Duration
9d 20h
Detected by Pingoru
Sep 13, 2026, 11:00 PM UTC

Affected components

Nango Cloud HealthNango Cloud Health

Update timeline

  1. investigating Sep 13, 2026, 06:00 PM UTC

    Some syncs may be failing to persist records. We are investigating the root cause.

  2. investigating Sep 13, 2026, 11:00 PM UTC

    Some syncs may be failing to persist records. We are investigating the root cause.

  3. resolved Sep 14, 2026, 03:02 AM UTC

    This issue has been resolved

  4. resolved Sep 23, 2026, 07:48 PM UTC

    Degraded record sync performance post-mortem Post-incident summary Date: 12–14 September 2026 Impact: Slower record syncs, with >1% failing on timeout and retrying. Most syncs completed successfully. Status: Resolved Summary Our records database experienced progressive performance degradation over a period of roughly two days. Customers experienced slower sync runs, and a subset (>1%) of syncs failed with timeout errors and had to retry. Most syncs continued to complete successfully throughout. Customers syncing large record volumes were significantly more likely to see failures, and some experienced repeated failures on specific syncs over a period of hours. Other parts of the system were not affected. No data was lost. Timeline Issue began: 12 September, 16:00 UTC Issue detected: 13 September, 23:05 UTC Mitigated: 14 September, 02:30 UTC Fixed: 17 September, 16:49 UTC Root cause A background job removes stored records for connections that have gone inactive. Because of a change made earlier this year, it was performing all of its work as a single, long-running database operation rather than in small batches. While it ran, the database could not perform its routine internal cleanup, which reclaims obsolete data and can only run once no active operation still depends on it. Obsolete rows accumulated, so every query had to read progressively more to return the same results. As read costs rose, write operations began exceeding their time limit and failing — and because failed syncs retry, the degradation compounded. Resolution We restored the batching behavior to the cleanup job, so it now processes a bounded amount of data at a time and releases the database between batches. The database's internal cleanup resumed immediately. Read performance returned to its normal baseline the same day and has remained stable since. Syncs that failed during the incident recovered automatically on their next run; no customer action was required. Next steps Monitoring: The condition behind this incident was visible on our health dashboards throughout, but we had no monitor configured to alert on it. We are closing this gap with new monitors.