# Google tests global GKE inference routing across 17,000 nodes

_Published Monday, September 21, 2026 at 1:09 PM EDT · AI, Infrastructure · Latest · Tier 2 — Notable_

![Google tests global GKE inference routing across 17,000 nodes — Primary](https://storage.googleapis.com/gweb-cloudblog-publish/images/07_-_Containers__Kubernetes_iY4YTLa.max-2600x2600.jpg)

Google says a multi-cluster GKE Inference Gateway test distributed foundation-model traffic across 17,000 compute nodes in three U.S. and European regions while retaining 99.5% of direct local-cluster throughput. Expanding from one cluster to three produced a near-linear throughput increase and a 99.9% request success rate under heavy concurrency.

The gateway used live key-value cache utilization, a measure of model-memory pressure, to spill requests from saturated regions to healthy ones behind one global address. The company tested one mixture-of-experts model with SGLang, so the results do not establish equivalent performance for other models or workloads.

## Sources

- [Cloud Blog](https://cloud.google.com/blog/products/containers-kubernetes/gpu-and-tpu-utilization-with-multi-cluster-gke-inference-gateway/)

---
Canonical: https://techandbusiness.org/newswire/oflZDk8F0a9_vpVdAdHs7q
Published: 2026-09-21T17:09:28.464Z
Story chronology: 2026-09-21T16:00:00.000Z
Retrieved: 2026-09-21T19:10:23.524Z
Publisher: Tech & Business (techandbusiness.org)
