“The DRA — Dynamic Resource Allocation — has now gone [generally available],” Macleod said. “The AI workloads are the real motivating factor, but it inspired a lot of really cool conversations with the ScheMD folks behind Slurm and their next round of adopters, who already run Kubernetes and are now adopting Slurm, but don’t want to learn how to run it on VMs or bare metal.”
GKE scalability evolves with AI-driven Kubernetes clusters
Google Kubernetes Engine scalability is moving from theory to practice as teams push clusters to unprecedented sizes — all while keeping costs and performance in balance. Enterprises are no longer questioning GKE scalability and are instead focusing on how to expose the appropriate controls in Kubernetes for measurable, predictable results.
As organizations stretch GKE to these new limits, they are finding that raw scalability is no longer the bottleneck. Instead, the real challenge is orchestrating heterogeneous workloads in a way that keeps infrastructure efficient, governed and fair across teams. This is exactly the gap new capabilities such as Dynamic Resource Allocation are addressing as artificial intelligence proliferates, according to Jago Macleod (pictured), director of engineering, Kubernetes, at Google Cloud.

Macleod, Gari Singh, product manager at Google Cloud, and Kate Holerhoff, senior industry analyst at RedMonk, spoke with theCUBE’s Savannah Peterson at the KubeCon + CloudNativeCon NA event, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed GKE scalability milestones in light of added node capabilities, control-plane resilience and AI-driven operations. (* Disclosure below.)
GKE scalability meets real-world constraints
AI adoption is raising the ceiling on cluster sizes while exposing very physical limits — power, cooling and specialized hardware mixes. As enterprises pair Kubernetes with AI workflows, they’re prioritizing standards and community-backed practices to keep agentic systems reliable and performant, according to Holerhoff.
At the upper end of scale, the practical drivers are massive-scale model training and the need to provision — and later shrink — fleets of graphics process units quickly. That pushes the control plane, scheduling and autoscaling to stay aware of health, updates and placement at extreme node counts, Singh noted.
“It’s usually in the massive training jobs, [the] massive AI jobs that need a lot of compute,” he said. “Typically, with nodes you can use the entire GPU … There’s typically a match of a node to a GPU. You’ll end up saying, ‘I need 130,000 GPUs to train whatever these massive models,’ … You need to quickly provision those up.”
Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the KubeCon + CloudNativeCon NA:
(* Disclosure: Google Cloud sponsored this segment of theCUBE. Neither Google Cloud nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)