CoreWeave expands full-stack AI cloud push as inference demand grows
The rise of AI-native cloud provider CoreWeave Inc. is part of the greater story emerging around operationalizing AI.
As enterprises shift from training to inference, neoclouds such as CoreWeave are providing cloud infrastructure tailor-made for AI. The company made waves by completing the industry’s first bring-up and validation of Nvidia Vera Rubin NVL72 on CoreWeave Cloud.
TheCUBE Research’s application development research shows that 86% of enterprises prioritize data unification over compute, reinforcing that AI performance depends on more than accelerators alone. That’s why CoreWeave’s support for Nvidia Vera Rubin, a unified, comprehensive platform built to support agentic AI, is so significant.
“CoreWeave’s evolution highlights a fundamental shift in how enterprises should think about AI infrastructure,” said Paul Nashawaty, principal analyst for theCUBE Research. “The challenge is no longer simply acquiring GPU capacity; it is building the integrated foundation needed to support AI applications throughout their lifecycle.”
SiliconANGLE Media’s livestreaming studio theCUBE will be on the ground at CoreWeave’s Fully Connected event in San Francisco from Sept. 30–Oct. 1, featuring insights from CoreWeave and its ecosystem partners. In advance of the event, we take a look at CoreWeave’s role in the changing AI marketplace. (* Disclosure below.)
AI infrastructure is now about having the full stack
Inference has become the new arena for companies looking to get a return on their AI investments and capital expenditures. As neoclouds and hyperscalers expand their offerings for AI workloads, CoreWeave is seeking to differentiate itself through expertise across the infrastructure stack.
“As agentic AI and inference workloads expand, enterprises will need infrastructure that supports scalable deployment, observability, security, governance and cost management across the application lifecycle,” Nashawaty said. “Neoclouds are positioned to play a growing role in that transition, but their long-term differentiation will depend on how effectively they connect infrastructure performance to measurable application outcomes.”
CoreWeave is positioning its purpose-built stack to deliver higher utilization, faster access to capacity and better token economics, advantages that customers might otherwise struggle to achieve alone. TheCUBE Research compares this value proposition to the early days of cloud, which reduced IT costs for customers and improved time to value.
Moving AI projects from experimentation into production remains a challenge across the enterprise market. Approximately 30% of organizations face an operational readiness gap between AI experimentation and production, while nearly 88% of AI pilots fail to reach production, according to theCUBE Research. By partnering with Nvidia Corp.’s full-stack platform, CoreWeave aims to narrow that gap.
“Neoclouds such as CoreWeave can close that gap by bringing together high-performance compute, data movement, infrastructure orchestration and the operational capabilities developers need to deploy and manage AI workloads reliably,” Nashawaty said. “CoreWeave’s validation of Nvidia Vera Rubin NVL72 shows the industry’s shift toward integrated, production-oriented AI systems rather than standalone infrastructure components.”
CoreWeave and Nvidia prepare for agentic era
Validating Vera Rubin is a major step for CoreWeave. TheCUBE’s analysts highlight cost compression as the platform’s most important commercial implication, with the new system offering one-tenth the cost per million tokens compared to Nvidia’s previous releases. However, improved token economics could expand the market rather than simply reduce spending.
“AI is no longer about isolated models,” said John Furrier, executive analyst for theCUBE Research. “It’s about systems — systems that bring together compute, networking, storage, software, data, security and operations into a unified platform capable of delivering real-world outcomes. The winners in this next phase won’t simply have access to AI; they’ll be the organizations that can operationalize it, scale it, govern it and continuously innovate around it.”
As AI infrastructure absorbs traditional general purpose IT functions, CoreWeave is positioning itself as an alternative to general-purpose cloud infrastructure. Full-stack platforms such as Vera Rubin are an essential part of reducing costs, since closer connections between models and data systems enable lower latency and faster inference feedback.
“There is so much demand in the market right now for the AI infrastructure that we provide,” said Jean English, chief marketing officer at CoreWeave. “We see this through clients coming to us who have tried something else and they didn’t get the reliability they needed. It’s bringing things up, bringing it up and validating it first to market as we did with Vera Rubin. All of that is an ecosystem that we serve and the ability for us to do that with the partnership that’s required is what we really see as a big differentiator.”
CoreWeave’s collaboration with Nvidia is just the beginning of the next stage of AI infrastructure, according to Chief Technology Officer Peter Salanki. He envisions a future in which inference is disaggregated into specialized pipeline stages. A smaller model would handle the query first, then larger models would take over for more complex tasks, giving users a balance of performance and cost.
The rise of agentic AI is also part of the equation, requiring an application loop with both GPUs and CPUs for secure code deployment.
“Rubin’s going to unlock the agentic era,” said Dion Harris, senior director of accelerated computing product marketing at Nvidia. “Agentic is basically where you’re taking not just a model and doing single-shot inference, but the model itself does planning. It has skills and subagents that … break down a problem and solve it to do real work. That’s why when we say [that] Vera Rubin was built for agents, it was built to enable that new workflow — the CPUs, the GPUs, the storage, all of that coming together — to unlock these agentic workflows.”
Neoclouds challenge hyperscalers for AI workloads
CoreWeave’s revenue backlog stood at approximately $104 billion as of June 30, representing contracted revenue that had not yet been recognized. The company also introduced new AI services this year, including an offering that enables enterprises to deploy AI agents that can improve autonomously using real-world data and its new Physical AI Field Engineering service.
The backlog and expanding portfolio reflect CoreWeave’s rapid growth as the company continues investing in infrastructure, capacity and services to address rising demand for AI workloads.
“Some interesting trends have developed that I think are going to make it very interesting at this [Fully Connected] event,” said theCUBE Research’s John Furrier, during an interview in advance of the conference. “One is a new term that has been kicked around in industry circles called ‘asset light,’ which means that you don’t have to spend billions and billions of dollars to get AI intelligence. That’s driving a lot of people to say, ‘Hey, I’ll just go to CoreWeave.’”
CoreWeave is betting that rising inference demand will strengthen the case for specialized AI cloud infrastructure. Salanki believes purpose-built AI infrastructure can offer advantages over traditional general-purpose cloud architectures, potentially giving neoclouds a distinct role as AI workloads expand.
CoreWeave’s value proposition encompasses more than accelerated compute; the company has redesigned facilities, racks and its software control plane. Vera Rubin racks reach up to 250 kilowatts per rack, making purpose-built facilities and technologies such as liquid cooling increasingly important to supporting high-density AI infrastructure. That requirement is a key part of CoreWeave’s infrastructure strategy.
“The agentic era demands a fundamentally different approach to infrastructure, one that keeps pace with workloads that reason continuously, scale unpredictably, and operate in production around the clock,” said Chen Goldberg, executive VP of product and engineering at CoreWeave. “What separates infrastructure that performs in a lab from infrastructure that performs in production is the depth of engineering underneath it.”
Stay tuned for the SiliconANGLE’s and theCUBE’s coverage of the Fully Connected event.
(* Disclosure: TheCUBE is a paid media partner for the Fully Connected event. Neither CoreWeave, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)
Image: SiliconANGLE/ChatGPT
A message from John Furrier, co-founder of SiliconANGLE:
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.