AI infrastructure demand is outrunning even the boldest supply chain strategies
AI infrastructure buildouts are moving so fast that plans made just months ago are already obsolete, forcing hardware makers to rewrite how they design, source and ship the systems powering the next generation of AI factories.
The shift from single-shot retrieval-augmented generation, or RAG, toward agentic AI has reshaped what enterprises need from their compute stacks almost overnight, according to Vik Malyala (pictured), chief business officer at Super Micro Computer Inc.
“Whatever we think is the pace at which the industry is moving, three months later, it is proving us wrong,” Malyala said. “Now we are talking about heavy use of agentic AI across the board, and we have a strong demand that is coming up not just on a GPU, but also on the CPU because of that.”
Malyala spoke with theCUBE’s Dave Vellante and John Furrier at the AMD Advancing AI event, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed AI infrastructure supply constraints, the AMD partnership and how agentic workloads are redrawing rack-scale design. (* Disclosure below.)
Supply constraints reshape AI infrastructure design choices
Component scarcity, particularly in memory, has forced Supermicro to work directly with suppliers such as Micron, Samsung and SK Hynix to secure allocation, Malyala explained. That constraint is reshaping customer expectations around delivery timelines as much as pricing.
“No matter how much planning that we have, we know for a fact the supply is a lot more constrained than ever before,” he said. “We set expectations with customers that it’s no longer possible to have a system shipped in a week or 10 days. It takes a longer time.”
Supermicro’s $60 billion order backlog reflects broadening adoption across banking, financial services and enterprise compute environments, not just hyperscale training clusters, Malyala noted. That demand is now shifting attention toward CPU-bound orchestration as agentic workloads multiply. Unlike RAG, which relies on a single inference loop, agentic AI spawns multiple sub-agents that reason, plan and take action across extended compute cycles — placing far greater pressure on the CPU to manage orchestration efficiently.
“Now you have the CPU, you have the GPU, you’re connecting — bringing AMD into the equation,” Malyala said. “Depending on what customers are trying to do, we can size it right.”
The co-design relationship with AMD is central to how Supermicro is preparing for that shift. As AMD moves toward higher-power GPU configurations and rack-scale systems such as Helios, Supermicro is redesigning its server portfolio — including its Hyper, CloudDC and Twin platforms — to use DC bus bar power delivery, matching the Helios architecture so data centers can deploy compute and GPU infrastructure without requiring separate power systems, Malyala said.
“Businesses need the cost of tokens to come down to a reasonable level,” Malyala said. “Then it will be adopted, and the business is profitable.”
Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the AMD Advancing AI event:
(* Disclosure: TheCUBE is a paid media partner for the AMD Advancing AI event. Neither AMD, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)
Photo: SiliconANGLE
A message from John Furrier, co-founder of SiliconANGLE:
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.