Lambda GPU Cloud incident

Degraded Performance in us-midwest-2

Lambda GPU Cloud is currently experiencing a major incident affecting Network and Virtual Machines, which began 19h ago. The vendor's full update timeline is below.

Started
Aug 22, 2026, 12:38 AM UTC
Resolved
Ongoing
Duration
● 18h 44m
Detected by Pingoru
Aug 22, 2026, 12:38 AM UTC

Affected components

NetworkVirtual Machines

Update timeline

  1. investigating Aug 03, 2026, 10:30 PM UTC

    Status: Investigating Severity Level: Medium We’re currently investigating a potential issue with our B200 cluster performance in us-midwest-2 impacting Managed Kubernetes, 1-Click Cluster, and Private Cloud customers. Our teams are actively working to identify the root cause and mitigate or resolve the issue. Thank you for your patience and understanding during this time. ----------------------------------------------------------------------------- In the event you’re unsure as to whether your service is impacted by this incident, please open a support ticket. We’ll get back to you as soon as possible. • https://support.lambdalabs.com/hc/en-us/requests/new To be automatically notified when the status of this incident changes, please click on the “Subscribe to updates” button Affected components Virtual Machines (Degraded performance) Network (Degraded performance)

  2. monitoring Aug 08, 2026, 02:10 AM UTC

    Status: Monitoring Severity Level: Medium We identified an issue within the InfiniBand fabric and implemented changes to address it. Performance has returned to expected levels based on telemetry and internal testing. We are continuing to monitor for stability. To be automatically notified when the status of this incident changes, please click on the “Subscribe to updates” button. Thank you for your patience and understanding during this time. Affected components Network (Degraded performance) Virtual Machines (Degraded performance)

  3. investigating Aug 11, 2026, 03:47 PM UTC

    Status: Investigating Severity Level: Medium We’re currently investigating a potential issue with our B200 cluster performance in us-midwest-2 impacting Managed Kubernetes, and 1-Click Cluster customers. Our teams are actively working to identify the root cause and mitigate or resolve the issue. Thank you for your patience and understanding during this time. Affected components Network (Degraded performance) Virtual Machines (Degraded performance)

  4. monitoring Aug 22, 2026, 12:38 AM UTC

    Status: Monitoring Severity Level: Medium Performance has returned to acceptable levels based on current telemetry and internal testing. We are continuing to monitor for stability and will provide further updates as they become available. ----------------------------------------------------------------------------- In the event you’re unsure as to whether your service is impacted by this incident, please open a support ticket. We’ll get back to you as soon as possible. • https://support.lambdalabs.com/hc/en-us/requests/new To be automatically notified when the status of this incident changes, please click on the “Subscribe to updates” button. Affected components Network (Operational) Virtual Machines (Operational)