Skip to content

UPDATED 11:00 EDT / SEPTEMBER 09 2026

AI

Lightbits set to release KV cache engine to boost GPU performance

Lightbits Labs Ltd. today announced the general availability of Inferra, a software engine designed to improve the economics and performance of artificial intelligence inference by moving key-value cache data beyond the limited high-bandwidth memory attached to graphics processing units.

The company, best known as the inventor of the NVMe over TCP storage protocol, is positioning Inferra for neocloud providers and enterprises running large language models with long context windows or many simultaneous sessions. Announced in March, the product manages KV cache across GPU high-bandwidth memory, dynamic random-access memory and Non-Volatile Memory Express storage, using predictive prefetching to put data near the GPU before it is needed.

KV caches hold the intermediate attention data that a model generates as it processes a prompt. Their size increases with the length of a conversation or document, consuming scarce GPU memory. When required cache data is unavailable, a system must retrieve it from slower memory or recompute it, leaving GPU resources idle and increasing response times.

“The data is always available for the compute to operate on,” said Ramesh Chettuvetty, senior vice president of product and business for AI solutions at Lightbits. “We do predictive prefetch, which essentially prevents this stall.”

Lightbits claims Inferra can cut time to first token by more than 100-fold in some long-context workloads, support context windows of more than 10 million tokens on commodity hardware and increase the density of concurrent sessions by more than 16 times. Those figures come from company benchmarks and extrapolations, and have not been independently verified.

Chettuvetty said Lightbits looks for signals throughout the inference stack to predict which data the GPU will need. He said the company has achieved cache hit rates of nearly 99.9% in most tested scenarios. Inferra also includes quality-of-service controls, encryption and isolation between tenants, and the ability to move cached data when a session migrates to another GPU cluster.

CPU-inspired

Arthur Rasmusson, director of AI architecture at Lightbits, compared the approach to processor techniques that were developed as CPU speeds began to outpace memory performance. “We’re doing something inspired by that,” he said. “The algorithms are obviously different when you’re dealing with LLM inference.”

The approach is intended to address two related problems: capacity and retrieval speed. Instead of loading an entire context into GPU memory, Inferra breaks it into smaller blocks and feeds the GPU what it needs just in time. That lets operators keep larger caches and fit more users on the same infrastructure.

“What we’re providing is sort of a larger bookcase in the sense that we have more ability to store long-term, but also fast retrieval,” Rasmusson said.

The largest benefits should come with retrieval-augmented generation, AI agents, lengthy prompts and heavily shared GPU services. Rasmusson acknowledged that the value would be more limited for organizations with oversized private GPU clusters, few users and contexts that already fit in available memory.

Disaggregating cache from the GPU introduces its own latency. Chettuvetty said Inferra addresses that delay by predicting demand far enough in advance to stage data before the GPU requests it. Faster networks and storage can narrow the required prediction window and reduce the risk of an incorrect forecast.

Lightbits said it has production pilots underway and is initially targeting neocloud providers, whose need to raise utilization and margins makes them faster adopters than hyperscalers. The company is demonstrating Inferra at the AI Infra Summit in Santa Clara next week.

Photo: Unsplash

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

 

About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

Send us a news tip

Send us a News Tip

  • This field is for validation purposes and should be left unchanged.
  • Max. file size: 244 MB.

Sign in

SIGN IN

Bio

Ethics statement

Extract the signal from the noise

Get SiliconANGLE updates and analysis.

Contact us

Partner with us

Contact us

Guest inquiry