AI
AI
AI
As artificial intelligence agents move from proof-of-concept tools to production systems, the cost of every generated token is becoming a direct business concern.
The shift is pushing infrastructure providers to focus not just on raw performance, but on efficiency, throughput and the economics of running agentic workloads continuously at scale. That pressure is reshaping how AI infrastructure is evaluated, according to Chen Goldberg (pictured), executive vice president of product and engineering at CoreWeave Inc.
“When your product is an AI agent, every token you generate has a cost and business impact,” Goldberg said. “Vera Rubin delivers 10 times better inference throughput per watt and one-tenth the cost per million tokens versus Blackwell. That is way more than a spec sheet improvement; that’s a technical leap.”
Goldberg spoke with theCUBE during “Scaling the Agentic Era With Nvidia Vera Rubin NVL72 on CoreWeave Cloud,” a virtual event hosted by theCUBE, SiliconANGLE Media’s livestreaming studio. Leaders from CoreWeave, Nvidia Corp. and Dell Technologies Inc. discussed how Vera Rubin is changing AI infrastructure, from end-to-end systems engineering to continuous operations and rack-scale validation. (* Disclosure below.)
Here’s theCUBE’s keynote session with Chen Goldberg:
Here are three insights you may have missed from the “Scaling the Agentic Era With Nvidia Vera Rubin NVL72 on CoreWeave Cloud” virtual event:
The infrastructure shift begins with the tighter coupling between models and the systems that feed them. As AI workloads become more interactive and stateful, performance depends on how reliably data, compute and inference feedback move through the system, according to Goldberg.
“That connection between the model and the infrastructure, it’s no longer just the GPUs,” Goldberg said. “We also need to make sure the data is getting there. We care about latency, and we care about … reliability because those are becoming mission-critical systems. This change in technology is blurring the line between training and inference.”
That shift is also changing how AI cloud platforms are built and managed. For AI agents, full-stack coordination is becoming central to making rack-scale systems operate as reliable production infrastructure, Goldberg noted.
“The first one, the brain of the system, [is] what we call Mission Control,” she said. “What we do there differently is that instead of thinking about compute, network and storage separately, we’re actually bringing those systems together and managing them as a whole.”
Here’s theCUBE’s complete video interview with Chen Goldberg:
The move from generative AI to reasoning systems, and now to AI agents, is changing the operational assumptions underlying AI infrastructure. Instead of serving a single model request, infrastructure must operate as a coordinated environment rather than a static deployment target, explained Harsh Banwait, director of product management at CoreWeave, and Dion Harris, senior director of accelerated computing product marketing at Nvidia Corp.
“Rubin’s going to unlock the agentic era,” Harris said. “Agentic is basically where you’re taking not just a model and doing single-shot inference, but the model itself does planning. It has skills and subagents that … break down a problem and solve it to do real work. To do that, you need a different type of infrastructure … and operations. That’s why when we say [that] Vera Rubin was built for agents, it was built to enable that new workflow — the CPUs, the GPUs, the storage, all of that coming together — to unlock these agentic workflows.”
At the application layer, that operating model shows up as feedback from production systems that can inform evaluations and new experiments. The infrastructure implication is that AI agents need platforms that can support that loop repeatedly, not just host an application after launch, pointed out Shawn Lewis, general manager of SaaS at CoreWeave, who talked to theCUBE along with Corey Sanders, senior vice president of product at CoreWeave.
“It can also … execute that entire loop,” Lewis said. “You can take an agent that’s running in production and say, ‘Find the areas that my agent’s not doing very well.’ The agent will look at those traces, cluster them together and tell you, ‘Your agent’s not doing very well on this type of problem.’ It’ll create an eval for you and run a new experiment.”
Supporting that kind of continuous AI operation also requires more precise workload orchestration. Because AI pipelines include multiple stages with varying latency and efficiency requirements, the value of AI infrastructure increasingly depends on whether applications can continuously and intelligently use the right resources, emphasized Peter Salanki, chief technology officer at CoreWeave.
“You don’t need Vera Rubin for every single thing,” he told theCUBE. “You can still use a different SKU … as an efficient part of your pipeline, be it speculative decoding … batch inference … or prefill. You can disassemble the pipeline into different pieces and put different things in there.”
Here’s theCUBE’s complete video interview with Harsh Banwait and Dion Harris:
As AI agents and larger models demand more memory, faster interconnects and denser systems, production readiness now depends on validating the rack as an integrated system before customers ever use it. That means testing power, cooling, networking, software and security together, according to Ihab Tarazi, senior vice president and chief technology officer of AI, compute and networking at Dell Technologies Inc., and Jacob Yundt, senior director of compute architecture at CoreWeave.
“The building block here is … that the rack is the new system now,” Tarazi said. “You’re not buying servers anymore; you’re buying those L11s. That system, while it looks easy, has thousands of components and dozens of pieces of software inside. We have a very extensive process to build, manufacture, design [and] test every component.”
CoreWeave’s first Vera Rubin rack shows how that validation model extends into live operations. For AI agents and large-scale model workloads, rack-level management now has to coordinate liquid cooling, power, sensors, observability and emergency response in real time. The company’s Racky rack manager and Valvey valve assembly are examples of that control layer, Yundt observed, along with Banwait and Zach Merendino, staff operations engineer at CoreWeave.
“Both of these technologies are going to allow us to control and observe our next generation of Nvidia Vera Rubin GPUs and CPUs,” Yundt said. “Racky truly is the brains of the operation. It connects all of the power, liquid cooling, sensors, management and observability into one centralized place that ties into both Mission Control and [CoreWeave Kubernetes Service]. This will allow us to scale to hundreds of thousands of GPUs, including Nvidia’s next-generation Vera Rubin.”
Rack-scale design is no longer simply an infrastructure milestone; it’s becoming the operating foundation for enterprise AI. By combining hardware density with software control, validation and orchestration, platforms such as CoreWeave Cloud are moving closer to the kind of production environment needed to support AI agents at scale, according to theCUBE Research’s John Furrier.
“AI is no longer about isolated models,” Furrier said. “It’s about systems — systems that bring together compute, networking, storage, software, data, security and operations into a unified platform capable of delivering real-world outcomes. The winners in this next phase won’t simply have access to AI; they’ll be the organizations that can operationalize it, scale it, govern it and continuously innovate around it.”
Here’s theCUBE’s complete video interview with Ihab Tarazi and Jacob Yundt:
To watch more of theCUBE’s coverage of the “Scaling the Agentic Era With Nvidia Vera Rubin NVL72 on CoreWeave Cloud” event, here’s our complete video playlist:
(* Disclosure: TheCUBE is a paid media partner for the “Scaling the Agentic Era” event. Neither CoreWeave, the sponsor of theCUBE’s coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.