Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
AI Classified by Officially
Today we’re releasing Olmo-core 3, a significant upgrade to our framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system.
Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency. It’s one of the core systems behind the next generation of Olmo, and part of our ongoing commitment to open up the tools and training infrastructure behind each new model.
Training large AI models takes a lot of compute, driving up costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs. MoE models offer a more efficient approach—they can contain many more learned components, or parameters, without requiring every input to use all of them. But the full model still has to be stored across GPU memory and updated during training, and directing inputs to the right experts – the specialized components within an MoE – across a cluster creates its own communication and coordination costs. As MoEs grow, those costs can erode much of the computational advantage of using only part of the model for each input.
Olmo-core 3 is built to close that gap. In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
The same infrastructure has been benchmarked at over one trillion total parameters.
Building a training stack around how MoEs actually work
Olmo-core has evolved with each generation of Olmo.
This is an extract. The publication continues at the source.
Read the original at the source: https://allenai.org/blog/olmocore3
Officially imported this from Allen Institute for AI’s own source and shows an extract. If you work there, claiming the profile and verifying the domain lets you choose to show the full text here.
Provenance
- Organization
- Allen Institute for AI — imported from official source
- Official source
- https://allenai.org/rss.xml RSS
- Imported
- October 01, 2026 16:00
- Versions
- 1 recorded
- Identity
https://allenai.org/blog/olmocore3