Olmo-core 3: Ai2 Speeds Up Mixture-of-Experts Training

Olmo-core 3: Ai2 Speeds Up Mixture-of-Experts Training

The Allen Institute for AI (Ai2), a Seattle-based AI research organization, released a new framework on Thursday for training large language models built on the mixture-of-experts design. The framework, called Olmo-core 3, is meant to push MoE training to the trillion-parameter scale while keeping compute use efficient and costs under control.

The code and related systems are available now on GitHub for developers and the open-source community.

Why mixture-of-experts models are hard to train

A dense model uses all of its parameters for every token it processes. A token is a small piece of text, often a word or part of a word, that a model reads or generates. Parameters are the internal values that determine how the model behaves.

A mixture-of-experts model works differently. It contains many specialist sub-networks, called experts, and sends each token to only a few of them. The model can hold far more parameters in total, but it only computes with a small share of them at any one step. That is the reason MoE designs have become popular for large models, including efficient open releases such as Qwen3.6-35B-A3B.

Training follows the same logic. Each token activates only part of the model, so the total compute drops. The catch is in memory and networking. The full model still has to fit in GPU memory somewhere, and the experts have to coordinate across the network during training. Both add cost.

Ai2 says Olmo-core 3 was designed to close the gap between dense and MoE training. In its setup, the pool of experts can grow from eight to 128 while each token still uses only four. On the same infrastructure, Ai2 says models can scale past 1 trillion parameters.

The benchmark numbers

Ai2 published one headline comparison. Training a 47-billion-parameter model on Nvidia B3000 GPUs, Olmo-core 3 processed 52,000 tokens per second. Nvidia's Megatron-core, an established framework for large MoE training, peaked at about 19,400 tokens per second in the same comparison. That works out to roughly 2.7 times the throughput.

These are Ai2's own figures. Independent tests on other model sizes and hardware will show how far the gain carries over.

How it saves memory

A white paper on the project describes three techniques that work together:

Expert parallelism. Experts are spread across several GPUs, so each card holds only part of the expert pool.

Layer splitting. The model's layers, the successive processing stages, are divided across groups of GPUs. Each GPU then keeps less of the model in memory.

A distributed optimizer. Training needs extra data, called optimizer state, to calculate and apply updates. Instead of keeping a full copy on every GPU, Olmo-core 3 spreads that state across many of them.

The combined effect is lower memory overhead as models grow, because no single device has to hold the whole model and its training state at once.

Ai2 also supports MXFP8, a number format that stores some values with fewer bits. That can cut both the amount of computation and the volume of data moved between GPUs.

Built for researchers, not just big labs

Ai2 frames the release as part of a broader goal: giving researchers the tools to build and train larger models. Trillion-parameter models are usually out of reach for anyone without government or enterprise-scale infrastructure.

The organization says Olmo-core 3 will also let researchers adapt MoE training to different hardware and experiment with routing, parallelism and other parts of the system. The stated aim is to grow a wider ecosystem around these methods.

Our Take

The interesting part of this release is not a new model. It is the plumbing. Most open-source attention goes to model weights, but training frameworks decide who can actually build large models in the first place. If Ai2's throughput numbers hold up outside its own benchmarks, a speedup of around 2.7 times over Megatron-core would be a meaningful cost reduction for academic groups and smaller labs.

This fits a broader pattern of open infrastructure work aimed at making large-scale AI less dependent on a few vendors, similar to Deepseek's open-source tooling for Huawei Ascend chips. Ai2's promise that researchers can adapt MoE training to different hardware points in a similar direction, though the benchmarks it published so far run on Nvidia GPUs.

Some caution is warranted. Faster training does not remove the need for large GPU clusters, and trillion-parameter runs will remain expensive. It is worth watching whether independent groups reproduce the reported throughput, whether the framework performs well on non-Nvidia hardware, and whether any notable open models trained with Olmo-core 3 appear in the coming months. That would be the clearest sign the tool is lowering the barrier, not just moving it.