UTHPC incident

UT HPC Cluster Rocket and Slurm upgrade

Minor Resolved View vendor source →

UTHPC experienced a minor incident on July 21, 2026, lasting —. The incident has been resolved; the full update timeline is below.

Started
Jul 21, 2026, 06:00 AM UTC
Resolved
Jul 21, 2026, 06:00 AM UTC
Duration
Detected by Pingoru
Jul 21, 2026, 06:00 AM UTC

Update timeline

  1. resolved Jul 21, 2026, 06:00 AM UTC

    Type: Maintenance Duration: 7 days, 10 hours and 6 minutes Affected Components: RStudio, Open OnDemand, rocket.hpc.ut.ee, Galaxy Jul 22, 13:40:21 GMT+0 - Identified - **Update:** Both login node updates have been completed successfully ahead of schedule. Login1 and Login2 are now fully available, and all temporary SSH connection restrictions have been lifted. Next, the Slurm upgrade on the 28th of July will follow. Until then, no interruptions in Rocket cluster computing are expected. Jul 28, 16:05:48 GMT+0 - Completed - Rocket cluster maintenace and SLURM update has been completed. Due to unfortunate circumstances, a subset of running jobs were impacted. Jul 17, 12:08:06 GMT+0 - Identified - Here is the summer maintenance schedule for the HPC Rocket Cluster. HPC cluster Rocket updates are scheduled for July 2026 to improve the cluster's performance and capabilities. **1\. Login Node Updates** We'll be performing system updates on both login nodes on the following dates: **\* Login1: July 21st** **\* Login2: July 28th** To minimize disruption, we'll close new SSH connections to the corresponding login node one week before each update, allowing existing connections to naturally expire. One of the login nodes will remain available at all times, so you won't experience any service downtime. **2\. Slurm Update** **On July 28th, starting at 15:00**, we'll be upgrading the Slurm version. Your running jobs won't be affected. However, during the update, submitting new jobs, SLURM commands like sacct, sacctmgr, and related tools will be unavailable. The process should take about two hours, but may run longer. The compute nodes will also receive OS, BIOS, and firmware upgrades. After July 28th, the compute nodes will be updated in a rolling fashion. This means some nodes will be temporarily drained until all updates are complete, which may result in longer queue times depending on cluster usage. The **Open OnDemand** service, as the UTHPC cluster portal, is also affected. Jul 21, 06:00:01 GMT+0 - Identified - Maintenance is now in progress