Linode incident

Emerging Service Issue - GPU Instances - All Regions

Notice Resolved View vendor source →

Linode experienced a notice incident on June 24, 2026 affecting US-East (Newark) and US-Central (Dallas) and 1 more component, lasting 1d. The incident has been resolved; the full update timeline is below.

Started
Jun 24, 2026, 06:53 PM UTC
Resolved
Jun 25, 2026, 07:29 PM UTC
Duration
1d
Detected by Pingoru
Jun 24, 2026, 06:53 PM UTC

Affected components

US-East (Newark)US-Central (Dallas)US-West (Fremont)US-Southeast (Atlanta)US-IAD (Washington)US-ORD (Chicago)CA-Central (Toronto)EU-West (London)EU-Central (Frankfurt)FR-PAR (Paris)

Update timeline

  1. investigating Jun 24, 2026, 06:53 PM UTC

    Our team is investigating an emerging service issue that is causing intermittent boot failures on GPU Instances in all Regions. We will share additional updates as we have more information.

  2. investigating Jun 24, 2026, 08:00 PM UTC

    We continue to investigate the issue that is causing intermittent boot failures on GPU Instances in all Regions. We will share additional updates as we have more information.

  3. monitoring Jun 24, 2026, 09:12 PM UTC

    At 20:32 UTC on June 24th, 2026 we have been able to correct the issue that is causing intermittent boot failures on GPU Instances in all Regions. We will be monitoring this to ensure that it remains stable. If you continue to experience problems, please open a Support ticket for assistance.

  4. resolved Jun 25, 2026, 07:29 PM UTC

    We haven’t observed any additional boot failures on GPU Instances in all regions, and will now consider this incident resolved. If you continue to experience problems, please open a Support ticket for assistance.

  5. postmortem Jun 26, 2026, 06:14 PM UTC

    On June 23, 2026, following a global software deployment, we identified an issue causing intermittent boot failures specifically for GPU Linodes. The impact was limited to instances where multiple GPU Linodes attempted to boot simultaneously. During the impact window customers could have experienced localized disruption or elevated error rates, particularly during automated node recycling or scaling events. During the investigation it was found that a recent software update created a conflict when multiple GPU servers tried to start up at the exact same time. The servers essentially blocked one another from loading, and our system did not automatically trigger a retry. This specific interaction only happens under heavy, simultaneous workloads, which is why it wasn't caught during our standard pre-release testing. In order to mitigate the issue, we successfully deployed a hotfix directly to all active GPU hosts across our fleet at 20:32 UTC on June 24th, 2026. The affected systems are currently operating normally as expected. In order to prevent this issue from happening in the future, we have developed and integrated a comprehensive fix into our upcoming software release scheduled to roll out globally over the next week. Additionally, we are actively prioritizing the procurement of dedicated GPU testing hardware for our development cloud to improve test coverage and ensure concurrent hardware workloads are fully simulated before future updates reach production. This summary provides an overview of our current understanding of the incident given the information available. Our investigation is ongoing and any information herein is subject to change.