Skip to content

UPDATED 20:26 EDT / AUGUST 26 2026

AI

Z.ai open-sources ‘Ox Alpha’ model as GLM-5.3-Flash

Z.ai Co. today released the code for GLM-5.3-Flash, a large language model that’s 10 times more cost-efficient than its predecessor.

The algorithm made its original debut last week under the codename Ox Alpha. LLM marketplace operator OpenRouter Inc. launched a free hosted version of Ox Alpha and didn’t disclose its developer, which drew a significant amount of industry attention. Users soon started speculating that Z.ai is the model’s creator.

GLM-5.3-Flash features a mixture of experts architecture with 320 billion parameters. It activates 18 billion parameters to answer prompts. Users requests can include a 1 million tokens worth of text, images and video while GLM-5.3-Flash’s responses contain up to 131,072 tokens.

The model features a different architecture than Z.ai’s earlier LLMs. One of the biggest changes is in its attention mechanism, a component that analyzes user prompts and extracts the most important details. It identifies those key details by breaking down each prompt into tokens and comparing them against each other.

Analyzing every single token in a lengthy prompt requires a significant amount of processing power. GLM-5.3-Flash reduces that hardware overhead with a technique called sparse attention. Instead of analyzing every single token in a prompt to find important details, the model reviews only the most relevant tokens.

Z.ai further reduced the LLM’s hardware footprint using a method called linear attention. Usually, doubling the size of a prompt quadruples the amount of memory that a model’s attention mechanism consumes. When linear attention is enabled, RAM usage only doubles.

One of the main reasons attention mechanisms are so memory-intensive is that they use an algorithm called a softmax function to interpret prompts. It turns numerical values generated by the host LLM into probabilities.  Linear attention, the technology Z.ai implemented in GLM-5.3-Flash, substitutes the softmax function with a more efficient algorithm.

The company says that the model costs 10 times less to run than its previous-generation LLM. Furthermore, it demonstrated strong performance across a set of popular artificial intelligence benchmarks.

Z.ai compared GLM-5.3-Flash against Claude Opus 4.8, GPT-5.6 Terra and Gemini 3.7 Flash. The former model achieved the highest score on GDPval-AA v2, an evaluation that measures LLMs’ ability to perform knowledge work. GLM-5.3-Flash also placed second on a benchmark called AutomationBench. It’s a test that assesses LLMs’ ability to complete tasks in cloud applications.

Z.ai trained GLM-5.3-Flash on a dataset with 30 trillion tokens. It used a technology called mHC to optimize the workflow.

When an LLM completes a training task, it receives a piece of data called a gradient. The data travels through the model’s artificial neuron layers and reconfigures them to improve their performance. The gradient sometimes becomes distorted along the way, which lowers its effectiveness. The mHC technology that Z.ai implemented in GLM-5.3-Flash lowers the risk of such technical issues. 

GLM-5.3-Flash’s weights are available on Hugging Face.

Image: Unsplash

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

 

About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

Send us a news tip

Send us a News Tip

  • This field is for validation purposes and should be left unchanged.

Sign in

SIGN IN

Bio

Ethics statement

Extract the signal from the noise

Get SiliconANGLE updates and analysis.

Contact us

Partner with us

Contact us

Guest inquiry