Google incident · via Google Maps Platform
Incident affecting Bare Metal Solution, Google Cloud NetApp Volumes and VMWare engine
Google experienced a major incident on July 15, 2026 affecting Bare Metal Solution and Google Cloud NetApp Volumes and 1 more component, lasting 12h 28m. The incident has been resolved; the full update timeline is below.
Affected components
Update timeline
- investigating Jul 16, 2026, 03:30 AM UTC
Description: GCVE: Our engineering teams are currently working with data center facility teams to address the rising temperatures. A shutdown has already occurred for some private clouds, and we expect that others will shut down as well. We recommend shutting down your VMware workload VMs for precaution so that recovery is safer. A shutdown of all systems started at 19:00 PDT [4:00 Local Time (GMT+2)]. We will notify you as soon as your Private Cloud has been powered back up and it is safe to restart the VMs. BMS: Some Bare Metal Solution servers are experiencing high temperatures and we are in the process of shutting them down. Customers may be unable to access their systems until we resolve the issue. Google Cloud NetApp Volumes Our engineering team has proactively shut down ONTAP clusters to protect the data stored on them. Customers may experience unavailability of Cloud NetApp Volumes until this issue is resolved. There is no estimated time for resolution. We will provide an update by Wednesday, 2026-07-15 21:30 PDT with current details. We apologize to all who are affected by the disruption. Customer symptoms: GCVE: Customers may have already observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and a recovery, we are shutting down all Private Clouds in this zone. BMS: Customers may have already lost connectivity to their BMS machines. As a protective measure for hardware integrity and a recovery, we are shutting down servers in this zone. Google Cloud NetApp Volumes Customers may have already lost connectivity to ONTAP clusters, or will lose connectivity to them shortly, as we proactively shut them down for their protection. Only Standard, Premium, and Extreme service levels are affected. Workaround: There are no workarounds available for this issue at this time. Customers with multi-regional deployments are advised to route traffic to an alternate site.
- investigating Jul 16, 2026, 04:17 AM UTC
Description: The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google has proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment. The datacenter operator team is working diligently to bring cooling and power back online so we can safely restart affected workloads. GCVE: Our engineering teams have shut down private clouds, and we expect that others will shut down as well. We will notify customers as soon as their private cloud has been powered back up and it is safe to restart the VMs. BMS: Some Bare Metal Solution servers are experiencing high temperatures and we are in the process of shutting them down. Customers may be unable to access their systems until we resolve the issue. Google Cloud NetApp Volumes Our engineering team has proactively shut down ONTAP clusters to protect the data stored on them. Customers may experience unavailability of Cloud NetApp Volumes until this issue is resolved. There is no estimated time for resolution. We will provide an update by Wednesday, 2026-07-15 22:30 PDT with current details. We apologize to all who are affected by the disruption. Customer symptoms: GCVE: Customers would have already observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and a recovery, we have shut down all private clouds in this zone. BMS: Customers may have already lost connectivity to their BMS machines. As a protective measure for hardware integrity and a recovery, we are shutting down some servers in this zone. Google Cloud NetApp Volumes Customers would have already lost connectivity to ONTAP clusters, as we proactively shut them down for their protection. Only Standard, Premium, and Extreme service levels are affected. Workaround: There are no workarounds available for this issue at this time. Customers with multi-regional deployments are advised to route traffic to an alternate site.
- investigating Jul 16, 2026, 04:41 AM UTC
Description: The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google has proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment. The datacenter operator team is working diligently to bring cooling and power back online so we can safely restart affected workloads. Some cooling has been restored, and we are working to restore affected workloads as available capacity allows. GCVE: Our engineering teams have shut down private clouds to protect these clouds from any damage. We will notify customers as soon as their private cloud has been powered back up and it is safe to restart the VMs. BMS: Some Bare Metal Solution servers experienced high temperatures and we proactively shut them down. Customers may have been unable to access their systems, but impacted systems have now been restored. Google Cloud NetApp Volumes: Our engineering team has proactively shut down ONTAP clusters to protect the data stored on them. Customers may experience unavailability of Cloud NetApp Volumes until this issue is resolved. There is no estimated time for resolution. We will provide an update by Wednesday, 2026-07-15 22:45 PDT with current details. We apologize to all who are affected by the disruption. Customer symptoms: GCVE: Customers would have already observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and recovery, we have shut down all private clouds in this zone. BMS: Customers may have lost connectivity to their BMS machines. As a protective measure for hardware integrity and recovery, we shut down some servers in this zone, but the impacted servers have been restarted and should be functioning normally. Google Cloud NetApp Volumes: Customers would have already lost connectivity to ONTAP clusters, as we proactively shut them down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels are affected. Workaround: There are no workarounds available for this issue at this time. Customers with multi-regional deployments are advised to route traffic to an alternate site.
- investigating Jul 16, 2026, 06:29 AM UTC
Description: The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment. Cooling and power have been restored at the site. The datacenter operator team is now working to safely restart affected workloads. GCVE: As a precautionary measure, our engineering teams had shut down private clouds to protect these clouds from any damage. Service restoration is still in progress. We will notify customers as soon as their private cloud has been powered back up and it is safe to restart the VMs. BMS: We are continuing to restore servers and monitor their health. Customers may be unable to access these systems until service restoration is complete. Google Cloud NetApp Volumes: Our engineering team had proactively shut down ONTAP clusters to protect the data stored on them. Customers may experience unavailability of Cloud NetApp Volumes until this issue is resolved. We will provide an update by Thursday, 2026-07-16 00:30 PDT with current details. Customer symptoms: GCVE: Customers would have already observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and recovery, we have shut down all private clouds in this zone. BMS: Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we shut down some servers in this zone. We are continuing to restore servers and monitor their health. Customers may be unable to access these systems until service restoration is complete. Google Cloud NetApp Volumes: Customers would have already lost connectivity to ONTAP clusters, as we proactively shut them down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels are affected. Workaround: There are no workarounds available for this issue at this time. However, customers with multi-regional deployments are advised to route traffic to an alternate site.
- investigating Jul 16, 2026, 07:39 AM UTC
Description: The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment. Cooling and power have been restored at the site. The datacenter operator team is now working to safely restart affected workloads. GCVE: As a precautionary measure, our engineering teams had shut down private clouds to protect these clouds from any damage. Network fabric has been restored, and Private Cloud restoration remains underway. We will notify affected customers once their Private Clouds are restored to allow for VM operational verification. BMS: We are continuing to restore servers and monitor their health. Customers may be unable to access these systems until service restoration is complete. Google Cloud NetApp Volumes: Our engineering team had proactively shut down ONTAP clusters to protect the data stored on them. Recovery is currently in progress. Customers may experience unavailability of Cloud NetApp Volumes until this issue is resolved. We will provide an update by Thursday, 2026-07-16 01:30 PDT with current details. Customer symptoms: GCVE: Customers would have already observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and recovery, we have shut down all private clouds in this zone. BMS: Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we shut down some servers in this zone. We are continuing to restore servers and monitor their health. Customers may be unable to access these systems until service restoration is complete. Google Cloud NetApp Volumes: Customers would have already lost connectivity to ONTAP clusters, as we proactively shut them down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels are affected. Workaround: There are no workarounds available for this issue at this time. However, customers with multi-regional deployments are advised to route traffic to an alternate site.
- investigating Jul 16, 2026, 08:32 AM UTC
Description: The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment. Cooling and power have been restored at the site. The datacenter operator team is now validating site stability. GCVE: Network fabric has been restored, and Private Cloud restoration remains underway. We are notifying the affected customers as their Private Clouds are restored to allow for VM operational verification. BMS: Server restoration is largely complete, with most systems now online and reporting a healthy status. We are actively working through the remaining validation efforts to achieve full recovery. Google Cloud NetApp Volumes: Service restoration is complete, and all affected servers are now operational. We continue to monitor these for full service recovery. If customers are still experiencing impact from this issue, please contact Support, and we will work with you to resolve any residual impact. We will provide an update by Thursday, 2026-07-16 02:30 PDT with current details. Customer symptoms: GCVE: Customers would have observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and recovery, we had shut down all Private Clouds in this zone. We are in the process of bringing all restored Private Clouds online. BMS: Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we had shut down some servers in this zone. We are continuing to restore servers and monitor their health. Google Cloud NetApp Volumes: Customers would have lost connectivity to ONTAP clusters, as they were proactively shut down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels were affected. Customers should now be able to access their volumes again. Workaround: There are no workarounds available for this issue at this time. However, customers with multi-regional deployments are advised to route traffic to an alternate site.
- investigating Jul 16, 2026, 09:21 AM UTC
Description: The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment. Cooling and power have been restored at the site. The datacenter operator team is now validating site stability. GCVE: Network fabric has been restored, and Private Cloud restoration remains underway. We are notifying the affected customers as their Private Clouds are restored to allow for VM operational verification. BMS: Server restoration is largely complete, with most systems now online and reporting a healthy status. We are actively working through the remaining validation efforts to achieve full recovery. Google Cloud NetApp Volumes: Service restoration is complete, and all affected servers are now operational. We continue to monitor these for full service recovery. If customers are still experiencing impact from this issue, please contact Support, and we will work with you to resolve any residual impact. We will provide an update by Thursday, 2026-07-16 03:30 PDT with current details. Customer symptoms: GCVE: Customers would have observed a loss of connectivity to their private cloud. As a protective measure for hardware integrity and recovery, we had shut down all Private Clouds in this zone. We are in the process of bringing all restored Private Clouds online. BMS: Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we had shut down some servers in this zone. We are continuing to restore servers and monitor their health. Google Cloud NetApp Volumes: Customers would have lost connectivity to ONTAP clusters, as they were proactively shut down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels were affected. Customers should now be able to access their volumes again. Workaround: None at this time.
- investigating Jul 16, 2026, 10:14 AM UTC
Description: The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment. Cooling and power have been restored at the site. The issue with GCVE and Google Cloud NetApp Volumes has been resolved for all affected users. We continue to recover the impact for BMS at the moment. GCVE [Mitigated]: Service restoration is complete across all the Private Clouds. Customers have been informed about this restoration and requested to verify that their VMs are fully operational. Google Cloud NetApp Volumes [Mitigated]: Service restoration is complete, and all affected servers are now operational. We continue to monitor these for full service recovery. BMS: Server restoration is largely complete, with most systems now online and reporting a healthy status. We are actively working through the remaining validation efforts to achieve full recovery. For Google Cloud NetApp Volumes and GCVE, if customers are still experiencing impact from this issue, please contact Support, and we will work with you to resolve any residual impact. We will provide the BMS restoration status update by Thursday, 2026-07-16 04:30 PDT with current details. Customer symptoms: The issue is now resolved for GCVE and Google Cloud NetApp Volumes. We continue to recover the impact for BMS at the moment. GCVE [Mitigated]: Customers would have observed a loss of connectivity to their private cloud. Google Cloud NetApp Volumes [Mitigated]: Customers would have lost connectivity to ONTAP clusters, as they were proactively shut down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels were affected. Customers should now be able to access their volumes again. BMS: Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we had shut down some servers in this zone. We are continuing to restore servers and monitor their health. Workaround: None at this time.
- investigating Jul 16, 2026, 11:39 AM UTC
Description: The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment. Cooling and power have been restored at the site. The issue with GCVE and Google Cloud NetApp Volumes has been resolved for all affected users. We continue to recover the impact for BMS at the moment. GCVE [Mitigated]: Service restoration is complete across all the Private Clouds. Customers have been informed about this restoration and requested to verify that their VMs are fully operational. Google Cloud NetApp Volumes [Mitigated]: Service restoration is complete, and all affected servers are now operational. We continue to monitor these for full service recovery. BMS: Server restoration is largely complete, with most systems now online and reporting a healthy status. We are actively working through the remaining validation efforts to achieve full recovery. For Google Cloud NetApp Volumes and GCVE, if customers are still experiencing impact from this issue, please contact Support, and we will work with you to resolve any residual impact. We will provide the BMS restoration status update by Thursday, 2026-07-16 05:30 PDT with current details. Customer symptoms: The issue is now resolved for GCVE and Google Cloud NetApp Volumes. We continue to recover the impact for BMS at the moment. GCVE [Mitigated]: Customers would have observed a loss of connectivity to their private cloud. Google Cloud NetApp Volumes [Mitigated]: Customers would have lost connectivity to ONTAP clusters, as they were proactively shut down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels were affected. Customers should now be able to access their volumes again. BMS: Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we had shut down some servers in this zone. We are continuing to restore servers and monitor their health. Workaround: None at this time.
- resolved Jul 16, 2026, 12:25 PM UTC
Description: The datacenter serving europe-west4-a for GCVE, BMS, and NetApp has experienced a power failure, which subsequently caused a cooling failure. Google proactively turned down workloads in order to protect customer data from any risks posed by running infrastructure in a high temperature environment. Cooling and power have been restored at the site. The issue with GCVE and Google Cloud NetApp Volumes has been resolved for all affected users as of Thursday, 2026-07-16 02:31 PDT and 01:00 PDT respectively. The issue with BMS is largely resolved and is believed to be affecting a very small number of customers. Our Engineering will continue to recover this residual impact directly with the impacted customers. If you have questions or are still impacted, please open a case with the Support Team and we will work with you until this issue is resolved. GCVE: Service restoration is complete across all the Private Clouds. Customers have been informed about this restoration and requested to verify that their VMs are fully operational. Google Cloud NetApp Volumes: Service restoration is complete, and all affected servers are now operational. BMS: Service restoration is largely complete. We continue working with a very small number of remaining affected customers to achieve full recovery. We thank you for your patience while we continue to fully resolve the issue. Customer symptoms: The issue is now resolved for GCVE and Google Cloud NetApp Volumes. The issue with BMS is largely resolved and our engineering will continue to recover this residual impact directly with the impacted customers. GCVE: Customers would have observed a loss of connectivity to their private cloud. Google Cloud NetApp Volumes: Customers would have lost connectivity to ONTAP clusters, as they were proactively shut down to protect these clusters from any damage. Only Standard, Premium, and Extreme service levels were affected. BMS: Customers may have lost connectivity to their servers. As a protective measure for hardware integrity and recovery, we had shut down some servers in this zone. Workaround: None at this time.
- resolved Jul 18, 2026, 12:22 AM UTC
Google cloud vmware engine impact: - Start: 15 July 2026 17:09 - End: 16 July 2026 2:33 - Duration: 9 Hours, 24 Minutes Google cloud netapp volumes impact: - Start: 15 July 2026 16:39 - End: 16 July 2026 01:10 - Duration: 8 Hours 31 Minutes Bare metal solution impact: - Start: 15 July 2026 18:37 - End: 16 July 2026 07:34 - Duration: 12 Hours 57 Minutes Summary On Wednesday, 15 July 2026, Google Cloud VMware Engine, Bare Metal Solution, and Google Cloud NetApp Volumes experienced service interruptions for a total duration of 14 hours, 55 minutes. We are taking immediate steps to ensure this doesn’t happen again. Preliminary Root Cause A regional data center hosting services in europe-west4-a experienced a loss of utility power and subsequent cooling capacity. The sequence of events leading to customer impact was as follows: - An electrical fault occurred on the utility grid upstream of the data center, disrupting the electrical distribution gear and cooling equipment. - The loss of cooling infrastructure resulted in ambient temperatures rising rapidly within the affected data halls. Host servers, storage clusters, and network switches were shut down to prevent equipment damages due to the extreme heat. - The high temperatures in the data hall resulted in disruption to customer workloads and control plane operations across the affected services. Google engineers have begun a full root cause analysis and we will provide additional information once it is available. Remediation Google engineering teams were alerted to the issue via automated temperature and hardware unreachability telemetry starting at 16:39 US/Pacific. Engineers collaborated with the third-party facility provider to safely restore primary utility power and cooling systems, returning ambient room temperatures to safe operational levels. With the environment stabilized, field support engineers and remote teams systematically booted and verified the server hosts, storage nodes, and network fabric in a controlled sequence. Automated recovery playbooks were executed to restore underlying network switches and routing infrastructure, allowing storage appliances and compute clusters to be brought back online and verified for health. The underlying infrastructure has been fully recovered, and normal operations have been successfully restored. Description of Impact On 15 July 2026 from 16:39 to 16 July 2026 07:34 US/Pacific, customers experienced service disruptions across several products in europe-west4-a zone: - Google Cloud VMware Engine: Customers lost connectivity to their private clouds and workloads as foundational hosts and network switches were powered off for thermal protection. - Bare Metal Solution: Customers experienced connectivity loss to their database servers and storage appliances as the underlying hardware was powered down. - Google Cloud NetApp Volumes: Customers in the STANDARD, PREMIUM, and EXTREME service levels were unable to access their storage volumes. Control plane operations, including the creation of new storage pools, volumes, and backups, experienced failures in the affected region.
- resolved Jul 25, 2026, 01:16 PM UTC
Incident Report Summary On Wednesday, 15 July 2026, Google Cloud VMware Engine (GCVE), Bare Metal Solution (BMS), and Google Cloud NetApp Volumes (GCNV) experienced service interruptions for a total duration of 14 hours, 55 minutes. The root cause for this outage is a 3ms voltage drop in the power feed from the utility provider and subsequent utility breaker protective action. This incident was mitigated by the data center provider rectifying the failed systems followed by the Google engineering team restoring the services. Root Cause A regional data center hosting services in europe-west4-a experienced an upstream voltage transient affecting both utility power feeds A and B. During this event, the utility breakers on both A and B feeds tripped and initiated the transfer to the back-up power source DRUPS (Diesel Rotary Uninterruptible Power Supply). The transition of side B to DRUPS system was successful without any power interruption. The DRUPS back-up power system for side A failed to take over the facility load due to electrical component failures. The DRUPS failure to take over load initiated automatic transfer of the affected 3 rows from feed A to redundant power feed B. Rows 1 and 2 transferred to power feed B successfully. Row 3 failed to transfer to power feed B and experienced complete power loss due to an overload protection breaker trip. The root cause of the failed transfer on the affected row was identified as load deployment discrepancy which is under further investigation by Google engineering team. This resulted in loss of both redundant power feeds to the single row. During the voltage transient event, the server data hall experienced an increased temperature due to a cooling system failure. The chiller controller dropped offline during the voltage transient event, failing to signal the chilled water distribution pumps to restart and ultimately causing the chiller system A to shut down. The redundant source was not available due to known ongoing construction work at the facility. This resulted in the data hall temperatures reaching 44°C and subsequent shutdown of the affected data hall machines. Google engineering teams initiated machine shutdown procedures for the remainder of reachable devices as part of the cooling emergency shutdown process. A timeline of events during the incident is provided below. DataCenter Events | Timestamp(PST) | Event Description | | :---- | :---- | | 07-15-2026 16:24 | An electrical fault occurred on the utility grid upstream of the data center, disrupting the electrical distribution. | | 07-15-2026 16:24 | DRUPS back-up power system for side A failed to take over the data hall load. | | 07-15-2026 16:24 | The chiller controller dropped offline during the voltage transient event causing distribution pumps to stop and unable to restart. | | 07-15-2026 16:37 | Rows 1 and 2 transferred to power feed B successfully. Row 3 failed to transfer to power feed B and lost power. | | 07-15-2026 18:29 | Data hall temperatures reached 44C and crossed the safe machine operating threshold. | | 07-15-2026 19:55 | Machine shutdown procedures implemented. | | 07-15-2026 21:05 | Cooling system fully recovered and temperatures in the data hall returned to normal operating range. | | 07-15-2026 21:24 | Notification about cooling recovery sent by provider | | 07-15-2026 21:46 | Notification about power recovery sent by provider | GCVE Timelines | Timestamp(PST) | Event Description | | :---- | :---- | | 07-15-2026 16:30 | GCVE service power monitoring alert received | | 07-15-2026 16:41 | First symptom detected, servers reporting redundancy power feed failure. | | 07-15-2026 17:05 | Switch temperature >60°C alerts triggered. | | 07-15-2026 17:09 | Start of user impact, first prober alert failure for north south traffic for customers. | | 07-15-2026 19:00 | Server / Network devices shutdown initiated. | | 07-15-2026 19:55 | All reachable server/network devices were shut down. | | 07-15-2026 21:24 | Cooling system fully recovered and temperatures in the datahall were returning to normal operating range. | | 07-15-2026 21:56 | Network restoration started with reachable hydra devices. | | 07-15-2026 22:40 | Onsite technician arrived, to recover the console servers. | | 07-15-2026 23:10 | Console server reboot/recovery complete. | | 07-16-2026 00:33 | Network recovery for all Placement Groups complete. | | 07-16-2026 02:36 | Incident mitigated, after recovering all the customer PCs are fully healthy. | This outage impacted 24 private clouds of 20 distinct customers in europe-west4-a region. Bare Metal Solutions (BMS) Timelines | Timestamp(PST) | Event Description | | :---- | :---- | | 07-15-2026 16:41 | First symptom detected, servers and storage reporting redundancy power feed failure | | 07-15-2026 17:52 | Automated monitoring detected rising temperatures | | 07-15-2026 18:37 | Four Netapp SAN storage nodes failed and were shut down due to overheating, resulting in storage availability issue | | 07-15-2026 18:37 | Customer impact started, first server shutdown detected | | 07-15-2026 18:59 | Reserve and buffer servers started to get powered off | | 07-15-2026 20:43 | Drop in the temperature has been detected | | 07-15-2026 20:52 | Four Netapp SAN storage nodes recovered, mitigating storage availability issue | | 07-15-2026 21:05 | The cooling system fully recovered and temperatures in the datahall were returning to normal operating range. | | 07-15-2026 23:23 | First customer server reboot started | | 07-16-2026 06:32 | Last customer server reboot started | | 07-16-2026 08:42 | Incident mitigated. All impacted customer servers were rebooted and confirmed in healthy state. | This outage impacted 9 distinct BMS customers in europe-west4 NetApp Timelines | Timestamp(PST) | Event Description | | :---- | :---- | | 07-15-2026 16:32 | Netapp facility remote monitoring alert on high temp and network switches switches failure received | | 07-15-2026 17:32 | All 6 clusters failed and were shut down automatically due to high temperature | | 07-15-2026 21:05 | Cooling capacity restored | | 07-15-2026 22:08 | Cluster recovery process initiated for all 6 clusters. 1st cluster recovered | | 07-15-2026 23:55 | All 6 Netapp Clusters and all customers recovered and validated as healthy & operational | Remediation and Prevention Google engineers were alerted to the infrastructure power loss on Wednesday, 15 July 2026 at 16:30 US/Pacific and subsequent increase in data hall temperatures at 17:44 US/Pacific via our automated monitoring and telemetry, and immediately began investigating. During the voltage transient event, the chillers maintained power but shut down because the distribution pumps failed to restart, resulting in no water flow. The pumps were manually switched from auto to hand mode to restore circulation and cooling. The facility team deployed an interim portable UPS to support the chiller controllers and prevent localized controller power loss. Upon confirmation of utility power restoration, the facility returned to utility power via automatic switching with the exception of two Remote Power Panels (RPPs), which required manual verification before restoration of power to affected server row. To provide full resiliency the facility team aligned a swing DRUPS unit to restore the redundant configuration on the A feed and is working on repairing the faulty DRUPS. To remediate the power loss at Row 3, once the electrical fault was confirmed to be fully isolated, teams reset the tripped breakers, successfully restoring both redundant power to the impacted server racks. Google is committed to preventing a repeat of this issue in the future and is completing the following actions: * Detailed investigation by the engineering team of the sequence of events that prevented transfer of load to DRUPS system and identified system improvements [ETA Aug 2026]. * Determine and resolve the cause of overload conditions that caused power loss to Row 3 [ETA Aug 2026]. * Perform investigation of chiller pump control system redundancy setup to ensure cooling system resiliency [ETA Sep 2026]. * Implement improvements to facility system monitoring alerts for power events and configure early alert thresholds for cooling excursion detection. [ETA Aug 2026]. * The GCVE engineering team is working to improve the automated triggering of shutdown during quickly evolving thermal runoff situations [ETA Oct 2026]. * The BMS engineering team is working with partner teams to update the incident classification and SLO definitions to improve personnel availability during major datacenter events [ETA Sep 2026]. * The GCNV engineering team is collaborate with the data center team to review incident handling SLA and runbooks, identify potential future occurrences and create playbooks to address/mitigate them, establish monthly joint emergency drill, and review if the Netapp cluster recovery process can be expedited [ETA Aug 2026]. Detailed Description of Impact On 15 July 2026 from 16:39 to 16 July 2026 07:34 US/Pacific, customers experienced service disruptions across several products in europe-west4-a zone: * Google Cloud VMware Engine: Customers lost connectivity to their private clouds and workloads as foundational hosts and network switches were powered off for thermal protection. * Bare Metal Solution: Customers experienced connectivity loss to their database servers and storage appliances as the underlying hardware was powered down. * Google Cloud NetApp Volumes: Customers in the STANDARD, PREMIUM, and EXTREME service levels were unable to access their storage volumes. Control plane operations, including the creation of new storage pools, volumes, and backups, experienced failures in the affected region.