Loading...
Loading...
A network configuration update caused inter-site communication failures for Google Cloud VMware Engine Stretched Clusters in australia-southeast2 and europe-west3 zones, triggering failover events. The issue was mitigated by rolling back the change and resetting network state.
High reliability impact due to inter-site communication failures and failover events in two zones. No direct cost, migration, quality, or compliance impact identified.
"modified": "2026-07-24T07:14:47+00:00", "created": "2026-07-24T07:10:18+00:00", "modified": "2026-07-24T07:14:47+00:00", "text": "## Incident Report\n## Summary\nOn Tuesday, 14 July 2026 10:00 PT, Google Cloud VMware Engine (GCVE) Stretched Cluster customers in the australia-southeast2 and europe-west3 zones experienced inter-site communication failures.\nThe disruption was traced to a network configuration update that introduced a conflict, causing inter-site communication failures, which triggered VMware Stretched Cluster failover events. Google engineers successfully mitigated the impact by rolling back the issue-causing configuration change and resetting the network state.\nWe sincerely apologize for the disruption this incident caused to your business. We know how much you rely on Google Cloud, and we regret the impact on your productivity. We are working to address the root cause and prevent this from occurring in the future.\n## Root Cause\nA network configuration update intended to prepare the cloud network infrastructure for new capabilities was deployed to the foundational network control plane. While the configuration payload itself was structurally valid, it exposed an implementation gap within the control plane's routing logic.\nThis logical gap caused the underlying network hosts to misconfigure program routing tables. As a result, traffic destined for the private IP address space used by GCVE Stretched Clusters was dropped. Because Stretched Clusters rely on this private address space to establish routing sessions for inter-zonal connectivity, the traffic drops severed communications between the active zones of the clusters, triggering VMware High Availability (HA) failovers.\nStandard routing health-checking protocols, such as Border Gateway Protocol (BGP) and Bidirectional Forwarding Detection (BFD), remained fully functional because their control plane sessions run on separate, unaffected address spaces. Because the control plane remained healthy, standard failover mechanisms failed to detect that the selective private IP range used for inter-zonal data tunneling was being dropped. Automated safeguards did not block the deployment because the configuration passed initial payload validations. The impact only manifested once the update began routing data traffic through the specific affected IP range.\n## Remediation and Prevention\nTo stabilize the environment, engineers identified the working network paths and deployed configuration changes to reroute traffic and restore connectivity. GCVE Stretched Cluster inter-zonal connectivity was completely restored for all supported locations on Tuesday, 14 July 2026 at 20:40 US/Pacific.\nGoogle is committed preventing a repeat of this issue in the future and is completing the following actions:\n* **Expanded Testing:** We are adding more detailed GCVE network setups to our existing testing environments. This allows us to automatically test future network updates against these configurations before they go live.\n* **Service-Level Data Path Failover:** We are implementing additional service-level data path failover mechanisms that actively probe the specific data-tunneling traffic space. This will ensure a path failover is triggered if the data plane itself is degraded even when the BGP control plane remains functional.\n* **Detailed Alerting**: We are adding faster, more specific alerts for connection issues between zones. This builds on our current platform monitoring to catch minor disruptions early and speed up our response.\n* **Improved Cluster Resilience:** We are fine-tuning the cluster's high-availability and storage settings. This makes virtual machines more resilient to short network drops, preventing them from restarting unnecessarily if the main site is still healthy.\n* **Workload Resilience Alignment (Shared Responsibility):** We are proactively reaching out to customers utilizing non-vSAN replicated virtual machine configurations within Stretched Clusters. Because these workloads are pinned to a single zone without active cross-site replication, they cannot survive inter-site network disruptions. We are ready to assist customers in auditing their storage policies, adjusting Stretched Cluster configurations, and planning the secondary zone capacity required to enable robust high-availability failovers.\n## Detailed Description of Impact\nOn Tuesday, 14 July 2026 from 10:00 to 20:40 US/Pacific, customers utilizing GCVE Stretched Clusters in the affected zones (australia-southeast2, and europe-west3) experienced inter-site communication failures. For some customers, depending on their architecture, this disruption led to a VMWare HA event causing VM restarts/movement across zones as designed, host disconnects, VSAN alarms and intermittent access to VMware Management (vCenter/NSX Manager) components.", "when": "2026-07-24T07:10:18+00:00" "created": "2026-07-24T07:10:18+00:00", "modified": "2026-07-24T07:14:47+00:00", "text": "## Incident Report\n## Summary\nOn Tuesday, 14 July 2026 10:00 PT, Google Cloud VMware Engine (GCVE) Stretched Cluster customers in the australia-southeast2 and europe-west3 zones experienced inter-site communication failures.\nThe disruption was traced to a network configuration update that introduced a conflict, causing inter-site communication failures, which triggered VMware Stretched Cluster failover events. Google engineers successfully mitigated the impact by rolling back the issue-causing configuration change and resetting the network state.\nWe sincerely apologize for the disruption this incident caused to your business. We know how much you rely on Google Cloud, and we regret the impact on your productivity. We are working to address the root cause and prevent this from occurring in the future.\n## Root Cause\nA network configuration update intended to prepare the cloud network infrastructure for new capabilities was deployed to the foundational network control plane. While the configuration payload itself was structurally valid, it exposed an implementation gap within the control plane's routing logic.\nThis logical gap caused the underlying network hosts to misconfigure program routing tables. As a result, traffic destined for the private IP address space used by GCVE Stretched Clusters was dropped. Because Stretched Clusters rely on this private address space to establish routing sessions for inter-zonal connectivity, the traffic drops severed communications between the active zones of the clusters, triggering VMware High Availability (HA) failovers.\nStandard routing health-checking protocols, such as Border Gateway Protocol (BGP) and Bidirectional Forwarding Detection (BFD), remained fully functional because their control plane sessions run on separate, unaffected address spaces. Because the control plane remained healthy, standard failover mechanisms failed to detect that the selective private IP range used for inter-zonal data tunneling was being dropped. Automated safeguards did not block the deployment because the configuration passed initial payload validations. The impact only manifested once the update began routing data traffic through the specific affected IP range.\n## Remediation and Prevention\nTo stabilize the environment, engineers identified the working network paths and deployed configuration changes to reroute traffic and restore connectivity. GCVE Stretched Cluster inter-zonal connectivity was completely restored for all supported locations on Tuesday, 14 July 2026 at 20:40 US/Pacific.\nGoogle is committed preventing a repeat of this issue in the future and is completing the following actions:\n* **Expanded Testing:** We are adding more detailed GCVE network setups to our existing testing environments. This allows us to automatically test future network updates against these configurations before they go live.\n* **Service-Level Data Path Failover:** We are implementing additional service-level data path failover mechanisms that actively probe the specific data-tunneling traffic space. This will ensure a path failover is triggered if the data plane itself is degraded even when the BGP control plane remains functional.\n* **Detailed Alerting**: We are adding faster, more specific alerts for connection issues between zones. This builds on our current platform monitoring to catch minor disruptions early and speed up our response.\n* **Improved Cluster Resilience:** We are fine-tuning the cluster's high-availability and storage settings. This makes virtual machines more resilient to short network drops, preventing them from restarting unnecessarily if the main site is still healthy.\n* **Workload Resilience Alignment (Shared Responsibility):** We are proactively reaching out to customers utilizing non-vSAN replicated virtual machine configurations within Stretched Clusters. Because these workloads are pinned to a single zone without active cross-site replication, they cannot survive inter-site network disruptions. We are ready to assist customers in auditing their storage policies, adjusting Stretched Cluster configurations, and planning the secondary zone capacity required to enable robust high-availability failovers.\n## Detailed Description of Impact\nOn Tuesday, 14 July 2026 from 10:00 to 20:40 US/Pacific, customers utilizing GCVE Stretched Clusters in the affected zones (australia-southeast2, and europe-west3) experienced inter-site communication failures. For some customers, depending on their architecture, this disruption led to a VMWare HA event causing VM restarts/movement across zones as designed, host disconnects, VSAN alarms and intermittent access to VMware Management (vCenter/NSX Manager) components.", "when": "2026-07-24T07:10:18+00:00" "modified": "2026-07-24T07:10:18+00:00",
"modified": "2026-07-24T07:14:47+00:00",
Check your models and usage for free. See the relevant official changes, estimated cost, and a practical next step.