UTHPC incident

GPU nodes on our Rocket cluster are currently unavailable

Major Resolved

UTHPC experienced a major incident on October 1, 2026 affecting rocket.hpc.ut.ee and Services (Galaxy) and 1 more component, lasting 1h 14m. The incident has been resolved; the full update timeline is below.

Started
Oct 01, 2026, 12:09 PM UTC
Resolved
Oct 01, 2026, 01:24 PM UTC
Duration
1h 14m
Detected by Pingoru
Oct 01, 2026, 12:09 PM UTC

Affected components

rocket.hpc.ut.eeServices (Galaxy)Services (Open OnDemand)

Update timeline

  1. investigating Oct 01, 2026, 12:09 PM UTC

    GPU nodes on our Rocket cluster are currently unavailable due to a recently disclosed vulnerability. All ongoing jobs on them will be killed as the nodes will undergo emergency maintenance. We apologize for any inconvenience and will inform you when the nodes are available for submission.

  2. identified Oct 01, 2026, 12:44 PM UTC

    GPU nodes on our Rocket cluster are currently unavailable due to a recently disclosed vulnerability. All ongoing jobs on these nodes will be stopped as the nodes undergo emergency maintenance. Once GPU node availability is restored, the affected jobs will be returned to the job queue. We apologize for any inconvenience and will inform you when the nodes are available again.

  3. resolved Oct 01, 2026, 01:24 PM UTC

    The emergency maintenance has been completed, and the GPU nodes on the Rocket cluster are available again. Affected jobs have been requeued and will run again as resources become available. Thank you for your patience and understanding.