Skip to content
theCUBE

UPDATED 13:02 EDT / OCTOBER 06 2026

AI agent memory is moving from GPUs to storage, Vast Data Inc. co-founder Alon Horev told theCUBE at CoreWeave Inc.'s Fully Connected event. AI

Vast uses tiered storage to ease AI agent memory demands

AI agent memory is creating new demands on infrastructure as agents run longer sessions and spread across the enterprise. Retaining that context and making it available when needed puts pressure on memory capacity and data movement.

Those demands extend beyond the context held during an individual interaction. Enterprise agents also need shared knowledge that persists across sessions, according to Alon Horev, (pictured) co-founder and chief technology officer of Vast.

“Memory for agents is a bit different,” Horev said. “First of all, there are multiple types of memory. There’s long-term memory where an agent can see past conversations and past interactions and look back and learn from its past experiences.”

Horev spoke with theCUBE Research’s John Furrier and Dave Vellante at Fully Connected 2026, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed AI agent memory, key value cache offloading and the shift toward data movement as the next bottleneck. (* Disclosure below.)

Why AI agent memory is moving beyond the GPU

The pressure shows up first in inference. Each long-running session holds its KV cache in graphics processing unit memory, and a session of half a million tokens can take up one-tenth to one-twentieth of a GPU’s memory, Horev explained.

“It’s also possible the agent would stop talking to the [large language model] because it’s compiling code, it’s testing software, or, as a human, I want to have a cup of coffee,” he said. “What you see is that if you could stretch that memory wall and basically offload those sessions to storage, you can avoid that repeat recalculation.”

Vast’s approach uses memory in tiers. GPU memory is used first, then central processing unit memory on the same machine, then persistent media that can hold petabytes of KV cache, with Nvidia Corp.’s Dynamo software orchestrating the process, Horev noted.

“You can move a session from one busy GPU to one less busy GPU and move KV cache either over the network or read it from Vast,” he said. “Once you look at inference as a distributed problem where you have the opportunity to use GPU memory, CPU memory and Vast across a fleet of machines, you have more optionality and you have more optimized scheduling.”

The stakes rise as companies deploy thousands of agents that handle sensitive data and act on customers’ behalf. Those enterprises need to record everything their agents do and retain it for a set period, which makes AI agent memory both a governance and performance asset, according to Horev. Vast has also launched a confidential computing service for sensitive workloads.

“These conversations that the agent is doing, it’s also gold,” he said. “It’s the same information that it can use for fine-tuning or training or creating purpose-built models.”

Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of Fully Connected 2026:

(* Disclosure: TheCUBE is a paid media partner for the Fully Connected event. Neither CoreWeave, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)

Photo: SiliconANGLE

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

 

About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

Send us a news tip

Send us a News Tip

  • This field is for validation purposes and should be left unchanged.
  • Max. file size: 244 MB.

Sign in

SIGN IN

Bio

Ethics statement

Extract the signal from the noise

Get SiliconANGLE updates and analysis.

Contact us

Partner with us

Contact us

Guest inquiry