Ai2 releases Olmo-core 3 to make developing large mixture-of-experts LLMs more efficient
Seattle-based artificial intelligence research firm Allen Institute for AI announced a development framework for large language models Thursday that significantly improves how mixture-of-experts large language models are trained.
The new framework, Olmo-core 3, allows MoE training to reach the trillion-parameter scale while keeping costs low by preserving computational efficiency.
Mixture-of-experts models operate differently from dense AI models: they split computation across specialist portions of the model each time a token is generated, whereas dense models activate the entire model. A MoE model may contain far more total parameters, the individual dials that tune model behavior, but it computes with only a small number of them as it generates each part of an answer.
Equivalently, training an MoE model allows activating only portions of it while learning each token, a small piece of text — often a word or a part of a word — that an AI reads and generates. This lowers the overall compute; however, the full model still must be stored across graphics processing unit memory and coordinating networking between experts during training still generates additional costs.
Ai2 said Olmo-core 3 was built to bridge the gap between dense models and MoE models, allowing the expert pool to grow from eight to 128 while still selecting only four experts per token. Using the same infrastructure, LLMs can scale to over one trillion parameters.
Benchmarking the next generation of efficient MoE training
In benchmarks, the company said Olmo-core 3 processed 52,000 tokens per second on Nvidia B3000 GPUs for a 47-billion parameter model. Compared to Nvidia Corp.’s Megatron-core training architecture, an established option for training large MoEs, which topped out around 19,400 tokens per second, this represented a jump of about 2.7 times the throughput.
In a whitepaper published about the project, Ai2 noted that the new architecture uses expert parallelism to spread experts across multiple GPUs, allowing each card to store only part of the full expert pool. It also splits the model’s layers, which are the successive stages that generate inputs, across groups of GPUs to reduce how much of the model each GPU needs to keep in memory. Finally, a distributed optimizer spreads the optimizer state, additional data used to calculate and apply updates during training, across multiple GPUs instead of storing full copies on every GPU.
Altogether, this reduces memory overhead as models scale because the entire model and its training state don’t need to be stored in memory all at once.
Additionally, Ai2 supports MXFP8, a number format for LLMs that represents some values with fewer bits. It can reduce the computation and amount of data moved between GPUs.
The new training architecture, and its higher efficiency, is part of the company’s vision to give researchers the tools to build and train larger models. Trillion-parameter AI models are often beyond the reach of those without access to state and enterprise infrastructure. Ai2 added that Olmo-core 3 would also allow researchers to adapt MoE training to different hardware, experiment with routing, parallelism and other parts of the system to build out a new ecosystem.
The project and related systems are currently available for developers and the open-source community on GitHub.
A message from John Furrier, co-founder of SiliconANGLE:
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.