pool.ntp.org incident

Database trouble

Notice Resolved View vendor source →

pool.ntp.org experienced a notice incident on September 11, 2026 affecting Management Portal and Public website and 1 more component, lasting 2h 22m. The incident has been resolved; the full update timeline is below.

Started
Sep 11, 2026, 07:28 PM UTC
Resolved
Sep 11, 2026, 09:51 PM UTC
Duration
2h 22m
Detected by Pingoru
Sep 11, 2026, 07:28 PM UTC

Affected components

Management PortalPublic websiteDNS updatesMonitoring System

Update timeline

  1. investigating Sep 11, 2026, 07:28 PM UTC

    The database cluster is crash looping and/or starting up very slowly after a crash. We're investigating.

  2. identified Sep 11, 2026, 07:49 PM UTC

    The symptoms and immediate cause have been identified and remedied. Working on restoring the database cluster to a healthy state.

  3. resolved Sep 11, 2026, 09:51 PM UTC

    This has been resolved. The cluster is healthy again and the underlying causes have been resolved.

  4. postmortem Sep 11, 2026, 10:01 PM UTC

    MySQL went down because our Ceph storage cluster filled up. That traces back to a Ceph security upgrade we'd just finished. The new version checks S3 upload signatures more strictly than before, and that exposed a bug in the S3 client our monitoring system uses: it doesn't sign one of the headers Ceph now requires. Its uploads got rejected, and instead of giving up it retried forever. Each upload is really several requests: the bulk data went through fine, but the small file that finalizes it kept getting rejected, so the data piled up without ever completing into anything usable. More than a terabyte of duplicate data accumulated in about half a day, more once you count replication overhead. A full Ceph cluster stops accepting writes from everything, and our database lives on the same cluster, so all three MySQL nodes crashed at the same time. Ceph was back as soon as we'd cleaned out the orphaned uploads. MySQL took longer: two of the three nodes had fallen too far behind to rejoin automatically, and their rejoin attempts took the primary down twice more before we fixed them by hand. The site was down a bit over two hours in total. Monitoring data from that window never got stored and can't be recovered. No other data was lost. The client library's bug is already fixed upstream, but the fix hasn't shipped in a release Mimir or Tempo use yet, so we moved that traffic to a connection where it signs everything correctly instead of waiting for it.