Olmo-core 3: Open, scalable infrastructure for large mixture-of-experts models
Olmo core 3: Open, scalable infrastructure for large mixture of experts models Ai2 has released Olmo core 3 , an updated framework for developing large language models with a re...
By Software Development Team
Olmo-core 3: Open, scalable infrastructure for large mixture-of-experts models
Ai2 has released Olmo-core 3, an updated framework for developing large language models with a redesigned open mixture-of-experts (MoE) training system.
The framework is designed to scale MoE training into the trillion-parameter range while maintaining computational efficiency. It supports the next generation of Olmo and makes the associated training infrastructure available for researchers and developers.
Scaling MoE models efficiently
Large AI models require substantial computing resources, which increases cost and energy use. Mixture-of-experts models can provide a more efficient alternative by containing many specialized components, or experts, while activating only a subset for each input.
However, the full model still needs to be stored across GPU memory and updated during training. Routing each token to the appropriate experts across a GPU cluster also introduces communication and coordination overhead. As MoE models grow, these costs can reduce the benefits of activating only part of the model for each token.
Olmo-core 3 is designed to address these challenges. In one benchmark, the expert pool increased from 8 to 128 while the system continued to select only four experts per token. The number of active parameters per token remained approximately constant at 3.2 billion, while total parameter capacity increased from 4.6 billion to 47 billion. Training throughput declined by less than 5%.
The same infrastructure has also been benchmarked with more than one trillion total parameters.
A training stack designed for MoEs
Olmo-core has developed alongside successive versions of Olmo. Earlier work on sparse models included OlmoE, which used 64 routed experts. Olmo 3, in contrast, used a dense architecture in which nearly the entire model was active for every token.
Olmo-core 3 adds a training system designed specifically for much larger MoE models. The earlier MoE implementation used fully sharded data parallelism (FSDP), gathering and resharding model weights for each small batch. The new system is based on distributed data parallelism (DDP). Experts remain resident on GPUs, and relevant data is routed to them without repeatedly gathering model weights.
In a preliminary test using eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack. The earlier implementation processed 19,400 tokens per second per GPU, making the new system approximately 2.7 times faster in that test.
Techniques for distributing and optimizing training
Olmo-core 3 combines several methods for distributing large MoE models across GPU clusters:
- Expert parallelism spreads experts across GPUs, so each GPU stores only part of the expert pool.
- Pipeline parallelism divides model layers across groups of GPUs, reducing the amount of the model that each GPU must keep in memory.
- A distributed optimizer distributes optimizer state, the additional data used to calculate and apply training updates, across GPUs instead of storing a complete copy on every GPU.
Together, these methods allow an MoE model and its training state to exceed the memory capacity of a single GPU without requiring every GPU to store the entire model.
The framework also includes optimizations for routing and expert computation:
- Rowwise expert parallelism places routed data directly into expert input buffers, reducing data rearrangement.
- GPU-resident routing keeps routing metadata on GPUs, allowing the CPU to queue work without waiting for metadata to be copied back.
- Grouped GEMM combines many small expert computations so GPUs can execute them more efficiently.
Olmo-core 3 supports MXFP8, a lower-precision numerical format that represents some values with fewer bits. Lower precision can reduce computation and the amount of data transferred between GPUs, provided the savings exceed the cost of converting between formats.
In a controlled benchmark using four NVIDIA B300 GPUs, with work distributed uniformly across experts, MXFP8 increased end-to-end training throughput by approximately 21% compared with BF16, the higher-precision baseline. Peak active memory declined from 103 GiB to 95 GiB. Most of the improvement came from feed-forward computation and data transfers between experts rather than from attention alone.
These optimizations affect one another. Faster computation can increase data-movement costs, while moving fewer bits may not improve performance if format conversion takes too long. Olmo-core 3 is designed to manage these trade-offs across the complete training process.
Scaling toward one trillion parameters and beyond
Ai2 benchmarked Olmo-core 3 across multiple configurations on NVIDIA B300 GPUs, including a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs. The highest observed throughput was 858 TFLOP/s/GPU, measuring useful model computation per second on each GPU.
These tests used random routing to evaluate system performance rather than the quality of a trained model.
Experiments with DeepEP v2, an alternative method for handling communication between experts across GPUs, reached a configuration with 2.38 trillion total parameters. This was a short-capacity test rather than a complete training run, so it demonstrates the scale supported by the infrastructure rather than sustained training performance.
The technical report also describes experiments that influenced the team’s approach to training and measurement:
- A score intended to encourage balanced routing could improve even when the actual workload became less balanced. This failure mode was termed token gerrymandering.
- Lowering experts’ learning rates because they processed fewer tokens did not improve results in the model family tested.
- GPU calculations took different amounts of time when processed values changed, even when matrix dimensions remained the same. Performance comparisons therefore require matching input values as well as matching shapes.
- Overlapping communication and computation on separate GPU streams did not always improve training speed. In some tests, it slowed end-to-end execution, showing that more overlap does not necessarily produce higher throughput.
The report documents these findings, along with the techniques that were tested and the approaches that were not adopted.
Infrastructure for the next Olmo generation
Olmo-core 3 is intended to serve as the foundation for the next generation of Olmo. That model will use an MoE architecture and is planned to be trained on Ai2’s largest dataset with its longest context window.
The new training stack extends beyond Ai2’s earlier MoE work while allowing the training process to adapt as models and hardware change. Researchers and developers can use Olmo-core 3 to train their own MoE models, adapt the framework to different hardware, and experiment with routing, parallelism, and other system components.
The framework, technical report, and implementation are available through the Olmo-core 3 project and its GitHub repository.