Olmo-core 3: Ai2 Speeds Up Mixture-of-Experts Training
The Allen Institute for AI (Ai2), a Seattle-based AI research organization, released a new framework on Thursday for training large language models built on the mixture-of-experts design. The framework, called Olmo-core 3, is meant to push MoE training to the trillion-parameter scale while keeping compute use efficient and costs under control.
The code and related systems are available now on GitHub for developers and the open-source community.
Why mixture-of-experts models are hard to train
A dense model uses all of its parameters for every token it processes. A token is a small piece of text, often a word or part of a word, that a model reads or generates. Parameters are the internal values that determine how the model behaves.
