Astronomer incident

Tasks remain in the queued state on some worker queues.

Major Resolved

Astronomer experienced a major incident on September 8, 2026, lasting 7h 35m. The incident has been resolved; the full update timeline is below.

Started
Sep 08, 2026, 12:33 PM UTC
Resolved
Sep 08, 2026, 08:09 PM UTC
Duration
7h 35m
Detected by Pingoru
Sep 08, 2026, 12:33 PM UTC

Update timeline

  1. identified Sep 08, 2026, 12:33 PM UTC

    We have identified the cause. On Astro, each worker queue is backed by a Kubernetes autoscaling resource whose name is built from a fixed platform prefix, the deployment release name, and the worker queue name. For a very small number of customers, where that combined name exceeds the Kubernetes 63-character limit, the autoscaling resource is rejected at creation. As a result, workers do not scale for the affected queue and tasks routed there remain in the queued state. This is why the limit can be reached even when the worker queue name itself looks short: the release name and platform prefix already consume a large portion of the 63 characters before the queue name is appended. Workaround: If affected, until the permanent fix is deployed, please keep worker queue name at or below 10 characters. We are preparing a permanent fix. Further updates to follow.

  2. monitoring Sep 08, 2026, 07:42 PM UTC

    A fix has been implemented, and we are monitoring the results.

  3. resolved Sep 08, 2026, 08:09 PM UTC

    The incident has been resolved.