Skip to content

UPDATED 16:34 EDT / SEPTEMBER 10 2026

INFRA

Chipmaker Positron nabs $875M to speed up inference with consumer-grade memory

Positron AI Inc., a developer of inference appliances that don’t include high-bandwidth memory, today announced that it has raised $875 million in funding.

NEA, Atreides Management, Valor Equity Partners, Andra Capital, SemiAnalysis Capital and Silicon Graphics Inc. and Netscape Inc. co-founder Jim Clark led the round. They were joined by more than a dozen institutional investors. Positron is now worth $5 billion, a fivefold increase from February.

Large language models frequently move data between the computing and memory circuits of the underlying graphics cards. The speed at which they do so is measured by a metric called memory bandwidth. The higher a chip’s memory bandwidth, the faster it can perform inference.

Server-grade graphics cards ship with HBM memory, the RAM variety with the highest memory bandwidth on the market. The supply of HBM is currently well below demand. Furthermore, there’s a shortage of the interconnects that are used to integrate HBM memory with graphics’ cards computing circuits.

Reno, Nevada-based Positron says that it has sidestepped the supply chain crunch. The company makes inference appliances that substitute HBM with LPDDR5X, a memory variety that is mainly used in smartphones.

LPDDR5X is more readily available than HBM and costs less. However, there’s a major catch: It has significantly less memory bandwidth, which usually translates into slower inference. Positron says that it has found a way to bridge the performance gap.

According to the company, most artificial intelligence accelerators use less than 30% of their HBM modules’ memory bandwidth. Positron’s inference appliances, in contrast, unlocks more than 90% of LPDDR5X’s throughput. The company says that the increased hardware utilization makes up for the difference in theoretical peak performance.

The company’s flagship system is called Titan. It ships with up to 18.4 terabytes of LPDDR5X that can provide 23.68 terabits per second of memory bandwidth. According to Positron, that large RAM pool enables a single Titan appliance to run an LLM with 32 trillion parameters and a context window of 10 billion tokens.

Titan performs inference calculations using a custom chip called Asimov. The processor is built around a systolic array, a set of identical computing modules that each include co-located memory. Asimov uses its co-located memory to store LLM weights, numerical values that play an important role in LLM output generation.

The chip’s systolic array is supported by modules optimized to run activation functions. Those are code snippets that determine which of an LLM’s neural networks should participate in an inference task. Additionally, Asimov includes central processing unit cores that function as a “programmable escape hatch.” They can take over tasks that the chip’s other modules aren’t optimized to perform.

Each Titan appliance includes up to eight Asimov chips. Customers can link together multiple systems into clusters with up to 16,384 accelerators.

Titan and Asimov aren’t yet in production. According to Positron, data from simulations indicates that a server rack powered by its silicon can process up to 26 times as many tokens per dollar than Nvidia Corp.’s Blackwell GB300 NVL72 appliance.

“Our focus now is to tape out Asimov, bring Titan to production, and scale manufacturing to meet the demand in front of us,” said Chief Executive Officer Mitesh Agrawal. “This financing gives us the resources to do exactly that.”

Positron expects to tape out Asimov at the end of the year using Taiwan Semiconductor Manufacturing Co.’s three-nanometer node. Mass production is set to follow suit in the second half of 2027. In conjunction, Positron will ramp up manufacturing of its Titan appliances. That effort will place an emphasis on securing LPDDR5X supply commitments from partners.

Image: Positron

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

 

About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

Send us a news tip

Send us a News Tip

  • This field is for validation purposes and should be left unchanged.
  • Max. file size: 244 MB.

Sign in

SIGN IN

Bio

Ethics statement

Extract the signal from the noise

Get SiliconANGLE updates and analysis.

Contact us

Partner with us

Contact us

Guest inquiry