Skip to content

UPDATED 18:38 EDT / SEPTEMBER 23 2026

INFRA

The AI factory is becoming the computer and it’s changing the semiconductor race

The next phase of artificial intelligence infrastructure will not be defined by a single graphics processing unit, chip architecture or model. Compute, memory, networking, packaging, power and software are converging into a new systems architecture with sovereignty emerging as a consequence of that shift.

The semiconductor industry is entering a new phase of the artificial intelligence buildout, and I think the market is still underestimating how profound the architectural change will be.

For the first several years of generative AI, the conversation centered on the accelerator. GPUs were the scarce resource, Nvidia Corp. became the defining company of the AI infrastructure cycle, and everyone wanted to know how many GPUs could be obtained and how quickly they could be deployed.

That was phase one.

What I’m seeing now is the center of gravity moving outward from the chip. 

The AI infrastructure discussion is rapidly becoming about systems: central processing units, GPUs and custom XPUs; memory and high-bandwidth memory; scale-up and scale-out networking; chiplets; advanced packaging; optics; cooling; software; electricity; and ultimately entire campuses operating as enormous computers.

The AI factory is becoming the computer.

And when that happens, the economics, competitive dynamics and even the geopolitical implications of semiconductors change with it.

I recently had an opportunity to hear several leaders I’ve spent considerable time interviewing over the years, including Advanced Micro Devices Inc.’ Chief Technology Officer Mark Papermaster, Broadcom Inc. semiconductor chief Charlie Kawwas, Samsung Electronics Co. Ltd.’s Paul Cho and OpenAI Group PBC hardware leader Richard Ho.  The discussion focused on what the semiconductor ecosystem may look like in the coming years.

My takeaway wasn’t about any one roadmap. It was that AI is forcing the industry to redesign the computer at the same time AI begins redesigning how the computer itself gets built.

From the GPU era to the heterogeneous AI factory

The notion that one processor architecture wins everything is increasingly hard to reconcile with where AI workloads are heading. Training, reasoning, inference, retrieval, agents and increasingly specialized enterprise workloads have radically different compute characteristics.

Papermaster made the case that broad-purpose CPUs and GPUs aren’t going away. Instead, modular architectures and chiplets allow semiconductor suppliers to create workload-specific variations while retaining common foundations. AMD’s multiple Venice designs are an example of that approach, including configurations increasingly tailored toward emerging agentic workloads.

At the other end of the spectrum are custom accelerators.

Kawwas framed XPUs as a market primarily available to a relatively small number of frontier AI companies operating at enormous scale. That’s an important distinction.

Custom silicon isn’t simply about designing a better chip. It requires enough workload volume, software-stack ownership and predictable demand to amortize the development cost. That’s why hyperscalers and frontier model companies are moving deeper into silicon.

Google LLC did it with tensor processing units. Amazon Web Services Inc. has Trainium and Inferentia. Microsoft has Maia. Meta Platforms Inc. ontinues investing in its internal accelerator strategy. And OpenAI is now pushing its own architecture through Jalapeño with Broadcom.

OpenAI describes the rationale almost perfectly. Richard Ho said Jalapeño was optimized around the “kernels, memory movement, networking and serving patterns” that matter for frontier models. 

That’s full-stack optimization.

The frontier of AI infrastructure is moving from buying chips to designing systems around workloads.

Memory moves to the first page

One of the strongest trends coming out of our analysis is that memory needs to move toward the beginning of the computer architecture process. Historically, architects could largely design the compute system and then attach an appropriate memory hierarchy.

AI flips that assumption. The sheer volume of parameters and data movement means that memory bandwidth, capacity, power consumption and physical proximity to compute increasingly determine system performance.

The discussion focused on co-designing logic, memory and packaging, including HBM qualification and 3D-integrated memory rather than treating these elements as separate purchasing decisions.

Cho recently described the shift succinctly: “Memory is no longer a supporting player in AI infrastructure, it is becoming the design center.”  I think that’s exactly right. The AI infrastructure race isn’t simply a compute race anymore.

It’s becoming a data-movement race.

Every time information travels across a board, rack or data center, it costs time and energy. That is why we’re seeing increasingly aggressive integration of compute, HBM, advanced packaging and high-speed interconnect.

It also explains why memory suppliers suddenly occupy such an important position in the AI value chain. If the first AI infrastructure bottleneck was accelerators, HBM and advanced packaging demonstrated that the bottleneck can move.

And it will move again.

The bottleneck clock keeps moving

This is a theme I’ve been thinking about for some time: there is no permanent AI bottleneck. There is a bottleneck clock.

GPUs become scarce. Capacity responds. Then HBM becomes scarce. Packaging becomes constrained. Networking gets stressed. Transformer and power capacity become the gating factor. Then cooling, land and permitting enter the equation. As one constraint gets solved, pressure moves somewhere else in the AI factory.

The semiconductor industry is working through advanced packaging and large substrates as increasingly critical constraints, while memory qualification, 3D integration and thermal management are all becoming architecture-level considerations. This is an important shift for the technology industry and for investors.

Looking only at GPU shipments tells you less and less about how quickly AI infrastructure can actually be brought online. The relevant question becomes: How fast can the entire supply chain deliver a functioning unit of intelligence production? That’s a very different metric.

Power becomes architecture

Then there is electricity. Papermaster emphasized power as one of the central constraints on future AI systems, which is why chiplets, 2.5D and 3D packaging, photonics, thermal management and hardware-software co-optimization are all becoming so important.

The scale is massive. The industry is moving beyond thinking about individual racks to planning clusters and campuses measured in gigawatts. The discussion contemplated training environments consuming multiple gigawatts and future campuses potentially moving toward the five- to 10-gigawatt range.

At that scale, calling these things “data centers” starts to obscure what’s happening. These are industrial systems. The traditional data center consumes electricity to run applications. An AI factory consumes electricity, data and silicon to manufacture tokens and intelligence.

That changes the economic model. Energy effectively becomes an input into intelligence production: Power → compute → tokens → intelligence → economic value. Performance per watt therefore becomes one of the defining metrics of the AI era. It’s economics. When your unit of computing begins consuming gigawatts, a few percentage points of efficiency become massive amounts of capital.

AI begins designing the next AI factory

But perhaps the most fascinating development is happening one level deeper.

AI is no longer simply the workload running on semiconductor infrastructure. AI is beginning to help design the semiconductor infrastructure itself.

OpenAI’s Jalapeño development is an early glimpse. The team moved from initial design into manufacturing tape-out at extraordinary speed, with AI models used to accelerate parts of the design and optimization process. OpenAI says the program compressed design-to-production to roughly nine months. Ho captured the impact beautifully: “The models are giving superpowers to our engineers.”

But he added an equally important qualification: engineers still drive the work and remain responsible for the final decisions. That’s the model I expect across technical disciplines.

AI doesn’t eliminate the semiconductor engineer. It expands the design space the engineer can explore. Architecture alternatives that were too time-consuming to evaluate become possible. Verification cycles compress. Software development accelerates. Physical-design optimization improves.

And this creates a remarkable feedback loop: AI designs better chips → better chips create better AI factories → better AI factories create better AI → better AI designs better chips. That compounding cycle may prove more consequential than any individual semiconductor process-node improvement.

We’ve spent decades thinking about Moore’s Law primarily as a manufacturing phenomenon. The next era may combine traditional Moore’s Law with something resembling engineering-velocity law. The organization capable of running more design experiments, validating them faster and integrating hardware and software more tightly gains an advantage even before the next transistor arrives.

The system becomes the new unit of value

That’s the larger shift I see. The next wave of performance gains is being absorbed into system-level innovation. 

Performance increasingly comes from combining: process technology, chiplets, custom silicon, HBM, advanced packaging, networking, optics, cooling, software and power engineering.

In other words, the next 10X improvement doesn’t necessarily come from one semiconductor breakthrough.

It can come from orchestrating the parts of the system better.

This is why Nvidia’s rack-scale strategy has been so significant. It’s why AMD is expanding from silicon toward complete AI systems. It’s why Broadcom’s custom silicon and networking portfolio is strategically important. It’s why memory suppliers such as Samsung, SK Hynix and Micron are moving closer to the center of architectural conversations.

And it explains why Dell, Supermicro, Hewlett Packard Enterprise and the systems ecosystem have a much larger opportunity than simply putting GPUs into servers. Someone has to industrialize the AI factory.

And then sovereignty enters the conversation

This brings me to a second-order consequence that I think will become increasingly important.

Once the AI factory becomes strategic industrial infrastructure, the question of who controls it becomes unavoidable. That’s where sovereign AI enters. The first generation of sovereignty conversations focused primarily on data residency.

Where is my data? Increasingly, that’s only the first question. If an AI factory depends on foreign silicon, foreign memory, one networking ecosystem, one cloud control plane, one model provider or externally controlled operations, those dependencies matter.

That doesn’t mean every country needs its own semiconductor fab. And it certainly doesn’t mean every enterprise should attempt to vertically integrate everything. Sovereignty is not self-sufficiency. It is understanding, controlling and managing critical dependencies.

The semiconductor architecture is already moving in this direction. Cho emphasized designing memory, second sources and qualified geographies into the architecture early rather than discovering those dependencies after deployment. That’s as much a sovereign AI principle as a semiconductor supply-chain principle.

The same applies to Ethernet optionality. The same applies to custom versus merchant silicon. The same applies to energy. The same applies to models.

A country can own thousands of GPUs and still lack meaningful control over its intelligence infrastructure. That’s why I would distinguish between GPU sovereignty and AI sovereignty.

They aren’t the same thing.

The race moves from chips to intelligence production

We’re heading toward a world where the competitive unit in AI is increasingly the entire intelligence-production system. Not the GPU. Not the large language model. Not the data center.  The system.

That is where AI factories and Sovereign AI intersect. One is the physical and computational system for manufacturing intelligence. The other asks who controls that system, its critical dependencies and the economics around it. And that’s why the semiconductor roadmap is starting to look like something much bigger than a semiconductor roadmap.

The first phase of AI was a race for models and GPUs. The next phase is becoming a race to build the world’s most efficient, scalable and resilient systems for producing intelligence. The AI factory is becoming the computer. And the battle over who can build it, operate it, improve it and control its critical dependencies is just getting started.

Image: SiliconANGLE

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

 

About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

Send us a news tip

Send us a News Tip

  • This field is for validation purposes and should be left unchanged.
  • Max. file size: 244 MB.

Sign in

SIGN IN

Bio

Ethics statement

Extract the signal from the noise

Get SiliconANGLE updates and analysis.

Contact us

Partner with us

Contact us

Guest inquiry