AI
AI
AI
Agentic AI is driving key technology providers to rethink the computing architecture required to run rapidly expanding autonomous systems.
In response to this challenge, two leading tech companies have recently unveiled a key milestone involving system-level validation for an entire rack-scale architecture. In June, CoreWeave Inc. and Nvidia Corp. announced the first bring-up and validation of Nvidia Vera Rubin NVL72 on CoreWeave Cloud.
The announcement involved a fundamentally different approach to infrastructure, designed to provide an environment where workloads reason continuously, scale unpredictably, and operate in production around the clock. This is testing the limits of data bandwidth, as described by Chen Goldberg (pictured), executive vice president of product and engineering at CoreWeave.

TheCUBE’s John Furrier and CoreWeave’s Chen Goldberg talked about the latest announcements during the event.
“Vera Rubin is not an incremental upgrade – 72 Rubin GPUs, 36 Vera CPUs, 260 terabytes per second of NVLink 6 bandwidth inside a single rack, which is more data bandwidth than is used by the entire global internet,” Goldberg said. “The world is shifting from asking AI questions to having AI actually do things continuously at scale without stopping to close your laptop, agents writing code, running experiments and executing multi-step reasoning loops. This is exactly what Vera Rubin was architected for.”
Goldberg spoke during “Scaling the Agentic Era With Nvidia Vera Rubin NVL72 on CoreWeave Cloud,” a virtual event hosted by theCUBE, SiliconANGLE Media’s livestreaming studio. Executives from CoreWeave, Nvidia and Dell Technologies Inc. spoke with theCUBE about the engineering and operational requirements needed to support production-scale inference and agentic AI workloads, as well as what it would take to build accelerated computing infrastructure for this next phase of AI. (* Disclosure below.)
CoreWeave’s announcement and theCUBE’s event highlighted the important role of Vera Rubin NVL72 in supporting large-scale inference, persistent reasoning sessions and production AI workloads that require more than raw GPU density. Nvidia’s ability to supply advanced chip architecture has allowed CoreWeave to provide multiple systems solutions, including liquid cooling, rack control, networking and secure multi-tenant operations.
These innovations include CoreWeave’s liquid cooling solution – Valvey – which monitors flow rate, temperature, pressure and leak detection in real time.
“It manages the liquid cooling system in a software-defined way,” Goldberg said. “We can control a single valve at a sub-second timescale, so if we detect any leak, as an example, we take action immediately.”

CoreWeave’s Peter Salanki spoke with theCUBE about emerging trends in computing architecture.
The latest release also included Racky, a new unified rack control appliance specifically designed for aggregating power, cooling and environmental sensors into a standardized management surface. This allows each Vera Rubin rack to be managed as a cloud resource rather than a custom one-off build, which provides system administrators with a bigger picture, according to Peter Salanki, chief technology officer of CoreWeave.
“It takes in telemetry from the GPUs themselves, takes in telemetry from the power systems, takes in telemetry from different leak sensors and from the building management system and allows us to tie all these things together,” Salanki told theCUBE. “It’s all deployed locally in the pod. It will interact with Valvey, it will interact with other systems upstream and downstream to make decisions.”
With multiple CPUs and GPUs in a single rack, communication becomes particularly important. CoreWeave’s announcement includes multi-rail and multi-plane networking, with support for both Nvidia Quantum-X800 InfiniBand and Nvidia Spectrum-X Ethernet with RDMA over Converged Ethernet RoCE.
“The genius of this rack scale system is that it allows you to scale memory, scale compute, scale all the fabrics so that it can talk to each other at full line rate so that GPU number 1 can talk to GPU number 72 at the exact same speed,” said Dion Harris, product leader at Nvidia. “This is what gives you a very consistent, reliable way to scale your workload across the entire rack. When you scale out … that’s where you start to leverage our Spectrum-X co-packaged optics network. That allows you to scale out efficiently across the racks so that as workloads scale across NVL72s, you can run those efficiently as well.”

CoreWeave’s Harshdeep Banwait (center) and Nvidia’s Dion Harris (right) spoke with theCUBE about the integration of processor technology in the latest offerings.
CoreWeave is also leveraging Nvidia BlueField-4 DPUs or data processing units to enable secure, multi-tenant AI cloud operations. The goal is faster data access and lower latency. BlueField-4 allows tenants to run workloads across the full Vera Rubin computing platform while preserving control and security.
“From our standpoint, the hardest thing to do was, when all of it comes together, how does it look and feel in practice?” said Harshdeep Banwait, director of product at CoreWeave. “We brought it all together, went through the validation flow to make sure that the Vera CPUs work with the Rubin chips, to work with ConnectX NICs, to work with the BlueField-4 DPUs. It was just making sure that all of these components talk to each other together and act as a system to essentially unlock that level of performance.”
To implement a rack-scale platform such as the Vera Rubin NVL72, CoreWeave drew from its partner ecosystem. Dell provided the architectural backbone for the platform through its high-performance PowerEdge XE9812 servers.
Dell’s involvement is based on a belief that as AI models expand at trillion-parameter scale and context windows encompass millions of tokens, compute density is going to grow in importance, according to Ihab Tarazi, senior vice president and chief technology officer of Dell.

Dell’s Ihab Tarazi and CoreWeave’s Jacob Yundt talked with theCUBE about inference performance and computing density.
“The density metric is going to become very important. How much can you squeeze out of the density and performance?” Tarazi said. “First of all, all the new models that really matter to people are trillion parameter models. So, they no longer fit for the most part. If you want the full performance and you want some of the use cases, they’re not going to fit on an 8-way GPU standard server. They really need those NVL72 GPU systems.”
CoreWeave’s validation of the Vera Rubin NVL72 also underscores the need for inference performance that can support agentic AI in production. According to theCUBE Research, the journey toward a meaningful return on investment from AI is entering a new phase. It has moved from model innovation to operationalizing inference at scale in key business processes.
“As we’ve seen, the inference market has grown exponentially over even just the last couple of years,” said Corey Sanders, senior vice president of product at CoreWeave. “It’s the opportunity for Vera Rubin to now play this really interesting role of both supporting this massive buildup of training while also supporting a huge opportunity for massive inferencing at a cost and performance that I think before would have been impossible. Brand new workloads are coming to life as part of it.”
These new workloads are being driven by increased adoption of agentic AI, and CoreWeave has taken a series of actions over the past year to build a cloud infrastructure that supports it. This has included the acquisition of the AI model development firm Weights & Biases Inc. in 2025.

CoreWeave’s Corey Sanders and Shawn Lewis spoke with theCUBE about new tools to drive agentic AI.
A number of new Weights & Biases agentic AI tools have been added to the CoreWeave platform since then.
“For the first time, there is an agent inside of Weights & Biases that helps those AI users train AI models and build AI applications,” said Shawn Lewis, founder and chief technology officer of Weights & Biases at CoreWeave. “There’s also a feature in Weights & Biases called W&B Launch that connects the Weights & Biases toolkit, which you can think of as a UI that users spend time analyzing data in, back to infrastructure. So it allows W&B users to launch and execute jobs on CoreWeave infrastructure from the UI. Now an agent can, instead of a human … launch experiments onto the infrastructure.”
The partnership between CoreWeave, Nvidia and Dell illustrates an important step in the evolution of software and hardware design. As CoreWeave’s Senior Director of Compute Architecture, Jacob Yundt noted in his interview with theCUBE, the company’s approach to building its platform is shaped by scale and changes in how the computer itself is defined.

An example of the Nvidia Vera Rubin Superchip was on display during the event.
“The NVL72 products have really changed the landscape,” Yundt explained. “They change how you approach engineering. The rack has this high-speed interconnect, this ultra-high bandwidth, ultra-low latency connection. All of the CPUs, all the GPUs, they’re all in one rack-level package. You’re no longer thinking of things as just an individual server. I have to think of them now as racks and then that has a cascading effect, I have all these racks that are working together. Now the rack is the computer.”
In redefining the rack, CoreWeave is also validating how innovation at the hardware and software layers is unlocking new levels of performance, efficiency and scale for AI-driven applications. This will power agentic AI and the systems needed to implement them throughout the enterprise.
“AI is no longer about isolated models,” said theCUBE’s John Furrier. “It’s about systems, systems that bring together compute, networking, storage, software, data, security, and operations into a unified platform capable of delivering real-world outcomes. Today’s discussions highlight an industry moving beyond proof of concept and into execution. The winners in this next phase won’t simply have access to AI, they’ll be the organizations that can operationalize it, scale it, govern it, and continuously innovate around it.”
To watch more of the “Scaling the Agentic Era” event, here’s our complete video playlist:
(* Disclosure: TheCUBE is a paid media partner for the “Scaling the Agentic Era” event. Neither CoreWeave, the sponsor of theCUBE’s coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.