AI infrastructure reliability moves beyond the rack
AI infrastructure reliability depends on how well compute, networking, power and cooling operate together.
Turning early hardware breakthroughs into repeatable deployments requires coordination across engineering, manufacturing and operations. Dell Technologies Inc.’s partnership with CoreWeave Inc. has helped it adapt successive generations of rack-scale systems to different facilities, according to Sarat Krishnan (pictured, right), director of PowerEdge AI architecture and systems development engineering at Dell.
“You have data centers that require cooling coming in from the top, cooling coming from the bottom,” he said. “You have power whips of various sizes. We have learned now to build modular rack-scale infrastructure that can quickly adapt to all of these needs.”
Krishnan and Jacob Yundt (left), vice president of engineering, compute architecture, at CoreWeave, spoke with theCUBE Research’s Dave Vellante and John Furrier at the Fully Connected event, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed joint engineering, manufacturing diagnostics, liquid cooling and data center coordination. (* Disclosure below.)
Moving AI infrastructure reliability upstream
For CoreWeave and Dell, design work begins months or years before new systems enter production. Firmware settings, mechanical changes and deployment requirements are worked through ahead of the first rack, Yundt explained.
“The partnership is everything from just really tight co-engineering of solutions to being embedded at the factory, like … working on making this product better, figuring out how we can scale it, figuring out how we can deploy it,” he said. “I think it’s mutually beneficial because we’re able to leverage Dell’s supply chain muscle [and] their engineering experience, and they’re able to leverage our deployment, our operations, and what we see operating this at scale.”
That collaboration also informs validation of the rack as an integrated production system. Dell uses operating feedback from CoreWeave to refine diagnostics across server and rack assembly, Krishnan explained.
“It’s really about pushing your risk to the left,” he said. “It becomes increasingly expensive to remediate and fix something as you go later on in the manufacturing process; [it] becomes extremely expensive if you have to replace something in the data center. What we have done over the last two years is shift our most complex diagnostics — the ones that are going to capture true hardware failures — further and further left.”
From rack management to data center coordination
Liquid cooling adds another dimension to AI infrastructure reliability. CoreWeave’s Racky rack manager brings power, cooling and environmental telemetry into a unified control interface.
“We work with CoreWeave to integrate Racky in our L11 factories,” Krishnan said. “So by the time it ships, it’s already tested, and we have done a lot of joint innovation. One of the things we’ve been inspired by our partnership with CoreWeave is to build better leak detection mechanisms. [Leaks] can be catastrophic.”
The operating challenge for Nvidia Corp.’s Vera Rubin systems now extends across compute trays, switches, data processing units and network fabrics. Coordinating those elements with facility power and liquid cooling requires management across the full infrastructure stack, Yundt explained.
“For Vera Rubin and beyond, I think what we’re seeing is that it’s no longer the rack [that] is the computer,” he said. “It’s ‘This row is the computer,’ or ‘This data hall is the computer,’ or ‘This data center is the computer.’ What that means is tight systems integration across the entire stack.”
Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the Fully Connected event:
(* Disclosure: TheCUBE is a paid media partner for the Fully Connected event. Neither CoreWeave, the sponsor of theCUBE’s coverage, nor other sponsors have editorial control over the content on theCUBE or SiliconANGLE.)
Photo: SiliconANGLE
A message from John Furrier, co-founder of SiliconANGLE:
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.