Allen Institute releases Olmo-core 3 to democratize trillion-parameter model training
The Seattle AI research firm unveiled a new training framework that dramatically accelerates mixture-of-experts language models while reducing computational costs, enabling researchers without massive infrastructure to build models at scale.

The Allen Institute for AI, a research organization based in Seattle, unveiled a training framework on Thursday designed to substantially enhance how mixture-of-experts language models are developed. Olmo-core 3, the new framework, enables MoE training to scale to the trillion-parameter range without sacrificing cost efficiency through improved computational performance.
Mixture-of-experts architectures function distinctly from conventional dense models. Rather than activating every component of the model when processing each token—a small unit of text, typically a word or subword—MoE systems route computation through specialized expert modules. Though these models contain substantially more total parameters, the individual adjustments that shape model behavior, they only engage a subset during each generation step, reducing overall computational demand.
During training, this selective activation approach similarly conserves resources by engaging only relevant model sections as the system learns from each token. However, the complete model architecture must still reside across graphics processing unit memory, and coordinating communication between experts introduces additional overhead costs.
Ai2 engineered Olmo-core 3 to narrow the efficiency gap between dense and MoE approaches. The framework permits the expert pool to expand from eight to 128 experts while maintaining selection of just four experts per token. Leveraging identical infrastructure, language models can now reach beyond one trillion parameters.
Performance gains in mixture-of-experts training
Testing results demonstrated that Olmo-core 3 achieved 52,000 tokens per second throughput on Nvidia B3000 GPUs when training a 47-billion parameter model. This substantially outpaced Nvidia Corp.'s Megatron-core, the established standard for large MoE training, which reached approximately 19,400 tokens per second—representing a 2.7-fold improvement in processing speed.
According to a technical whitepaper, the architecture employs expert parallelism to distribute experts across multiple GPUs, ensuring each processor stores only a portion of the complete expert collection. The system also partitions model layers—the sequential stages responsible for generating intermediate representations—across GPU clusters to minimize individual GPU memory requirements. A distributed optimizer further alleviates memory strain by spreading optimizer state, the supplementary data needed for computing and implementing training updates, across GPUs rather than duplicating it on every device.
These design choices collectively diminish memory consumption as models grow larger, since the entire model and its training state need not occupy memory simultaneously.
The framework additionally incorporates MXFP8, a numerical representation format that encodes certain values using fewer bits. This capability reduces both computational load and data transfer volume between GPUs.
Ai2 positioned the training architecture and its efficiency gains as central to its mission of equipping researchers with resources for constructing and training increasingly capable models. Trillion-parameter systems typically remain inaccessible to those lacking connections to major institutional or commercial computing infrastructure. According to the institute, Olmo-core 3 enables researchers to customize MoE training for diverse hardware configurations, investigate routing strategies, parallelism approaches and other system components, thereby fostering a broader ecosystem of model development.
The framework and supporting systems are presently accessible to developers and the open-source community through GitHub.


