NetApp and Nvidia rethink storage for AI factories
Storage architecture is being rewritten for artificial intelligence factories. Traditional enterprise storage was designed around workloads that scaled in relatively predictable ways. AI changes that equation by combining heavy data movement with transactional metadata activity, often on shared infrastructure.
NetApp Inc. is addressing that pressure through its work with Nvidia Corp., a partnership that dates back more than a decade and now includes co-engineering around AI infrastructure. NetApp’s Novus architecture separates data and metadata functions so each can scale independently based on different workload demands while keeping GPU resources supplied with data, according to Arindam Banerjee (pictured, right), chief platform and technology officer of NetApp.
“What happens is different types of workloads like transactional workloads for metadata and heavy sequential workloads for, say, checkpoints happen at the same time,” Banerjee said. “Previous architectures did one or the other very well. They never were able to combine it and do both at the same time. This is what AI factories is pushing us, the workloads are pushing us to do both at the same time in terms of data.”
Banerjee and Jason Hardy (left), vice president of storage technology at Nvidia, spoke with Christophe Bertrand and Rebecca Knight at NetApp INSIGHT, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed how AI factories and agents are forcing storage systems to scale and operate differently. (* Disclosure below.)
Storage architecture shifts with AI scale
The separation of metadata from data management is intended to prevent smaller transactional operations from competing with large data transfers for the same resources. That matters because storage delays can leave costly GPUs underused even while they continue consuming power, making efficiency a larger part of the AI scaling equation, Banerjee explained.
“What we can do is scale each one of them independently on different axes,” he said. “If the metadata and the data shared the same resources, the same media, the same network, it resulted in an inefficient system. Because the data operations would be queued behind the metadata operations, which are small and transactional in nature. That not only left your GPUs underutilized, they were consuming power. To bring efficiencies back, to drive maximum out of our ecosystems and keep the GPUs fed, we needed to make this architectural shift.”
AI infrastructure also has to scale without forcing enterprises to redesign the system every time a new use case comes online. The goal is greater flexibility across fine-tuning, inference, capacity and AI-ready data as production workloads change and demand shifts across different parts of the infrastructure, Hardy emphasized.
“I think it all comes down to, one, you want to be able to design a system that allows you to scale over time,” he said. “When you step up … they’re moving into this phase 2 where we’re now past trying it and now it’s really pushing into being mainstream production. AI agents are really starting to show up now. Inferencing for enterprise scale is happening.”
Agents raise the concurrency bar
Agentic workloads bring a different kind of pressure because thousands of agents may need to access data at the same time. Their permissions may also be temporary and narrowly scoped, increasing the volume of metadata operations alongside the throughput demands already placed on storage systems, Banerjee noted.
“I think we talked about throughput. That’s important,” he said. “Another very important thing is concurrency. If you’re thinking agents coming and hitting your data, thousands of them coming at the same time, authorizing themselves … all this happening at the same time brings a very different set of concurrency requirements to your systems apart from the throughput. The metadata access through the concurrent agents is what is really redefining how we are going to build this thing for the future.”
That operational shift also changes who interacts with the storage layer. AI teams may need to provision and consume infrastructure without becoming storage specialists, making APIs, SDKs and agent-friendly controls increasingly important as these environments move into production and become part of everyday enterprise operations, Banerjee explained.
“Remove the complexity,” he said. “These are AI teams, they are not storage engineering teams. They should be able to consume storage through an API, code to an API, code to an SDK, for example. Or in the future, agents may be doing that work. It has to be really API-driven, agent-friendly consumption. That’s why we are designing a new control plane that allows the consumption of Novus through the APIs.”
Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of NetApp INSIGHT:
(* Disclosure: TheCUBE is a paid media partner for NetApp INSIGHT. Neither NetApp, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)
Photo: SiliconANGLE
A message from John Furrier, co-founder of SiliconANGLE:
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.