Complete AI Training

AI news ·

Ai2 releases Olmo-core 3, an open training framework for trillion-parameter mixture-of-experts models

Ai2 released Olmo-core 3, an open-source training framework that scales mixture-of-experts models to 1.2 trillion parameters while keeping throughput loss under 5% as expert count grows.

Share

Ai2 released Olmo-core 3 on October 1, 2026, a redesigned open-source training framework for large mixture-of-experts models. The system scales MoE training into the trillion-parameter range while keeping computational costs nearly flat-a 47-billion-parameter model with 128 experts lost less than 5% throughput compared to an 8-expert, 4.6-billion-parameter configuration.

The upgrade targets a core bottleneck in sparse model training. MoEs can hold many more parameters than dense models without activating all of them for every token. But storing the full model across GPU memory and routing data to the right experts creates coordination overhead. As expert pools grow, that overhead can erase the efficiency gains of using only part of the model per input.

"Olmo-core 3 is built to close that gap," the team said. In one benchmark, active parameters per token stayed roughly fixed at 3.2 billion while total parameter capacity grew from 4.6 billion to 47 billion. Throughput fell by less than 5%.

Switching from FSDP to a DDP-based architecture

Earlier Olmo-core MoE implementations used fully sharded data parallelism, which gathered and resharded model weights for each small training batch. Olmo-core 3 moves to a distributed data parallelism system that keeps experts resident on GPUs and routes data to them instead. This avoids repeated weight gathering across the cluster.

In a preliminary test on eight NVIDIA B300 GPUs, the new stack processed 52,000 tokens per second per GPU on a 47-billion-parameter MoE. The earlier FSDP-based implementation managed 19,400 tokens per second-roughly 2.7× lower throughput.

Parallelism, routing, and precision optimizations

The framework combines three distribution techniques. Expert parallelism spreads experts across GPUs so each device stores only part of the pool. Pipeline parallelism splits model layers across GPU groups to reduce per-device memory. A distributed optimizer spreads optimizer state across GPUs rather than replicating it.

Routing efficiency gets several targeted improvements. Rowwise expert parallelism places routed data directly into expert input buffers, cutting rearrangement work. GPU-resident routing keeps metadata on the GPUs so the CPU can queue work without waiting for data copies. Grouped GEMM combines small expert computations for more efficient GPU execution.

Olmo-core 3 also supports MXFP8, a lower-precision number format. In a controlled benchmark on four B300 GPUs, enabling MXFP8 where it helped most delivered about 21% higher training throughput than the BF16 baseline. Peak active memory dropped from 103 GiB to 95 GiB. Most gains came from feed-forward computation and data movement between experts, not attention alone.

Trillion-parameter scale and experimental findings

The team benchmarked a 1.2-trillion-parameter model with 58.36 billion active parameters per token across 512 GPUs, reaching 858 TFLOP/s/GPU. Experiments with DeepEP v2, an alternative expert communication system, reached a configuration with 2.38 trillion total parameters in a short-capacity test.

The accompanying technical report documents several counterintuitive findings. A routing balance score improved even as actual workload distribution worsened-a phenomenon the researchers call token gerrymandering. Lowering expert learning rates because they process fewer tokens did not improve results in the tested model family. GPU computation times varied even with identical matrix dimensions when input values changed, meaning performance comparisons need matching values, not just matching shapes. Overlapping communication and computation on separate GPU streams sometimes slowed end-to-end execution rather than speeding it up.

Why this matters for IT and research professionals

Olmo-core 3 is fully open. Researchers and developers can use it to train their own MoEs, adapt it to different hardware, and experiment with routing strategies and parallelism configurations. The framework gives direct control over the trade-offs between computation speed, data movement, and memory use that define large-scale training economics. For teams working with sparse architectures or evaluating infrastructure for next-generation models, the open stack removes a barrier that has kept advanced MoE training concentrated in a handful of well-funded labs. The code is available on GitHub, and the technical report details the approaches tested and the ones the team chose not to adopt.

Share