Article

Olmo-core 3: What Changed and How to Plan a MoE Training Migration

Ai2’s Olmo-core 3 adds a MoE training stack. Check benchmark limits, PyTorch extras, cluster topology and checkpoint behavior before migrating.

Editorial illustration for Olmo-core 3: What Changed and How to Plan a MoE Training Migration: an arrow connects an old service to its replacement. Not documentary evidence.

Ai2 released Olmo-core 3 on October 1, 2026, upgrading its open training framework for labs building mixture-of-experts (MoE) models. The release includes the OLMoDDP training stack and reference scripts. It is training infrastructure, not a release of trained trillion-parameter model weights. For existing users, the immediate task is to qualify dependencies, checkpoint continuation and cluster performance before moving a training run.

The announcement appeared at 15:01 UTC, followed by GitHub’s v3.0.0 release at 15:17 UTC. The 3.0.0 package is also available on PyPI.

What the new training path changes

Its MoE path keeps experts resident on GPUs and sends token data to them, reducing repeated weight gathering. Ai2 contrasts this with its earlier fully sharded data-parallel implementation. Expert parallelism distributes experts across GPUs; pipeline parallelism distributes layers; a distributed optimizer divides the update state. Together, these address memory and communication costs that sparse activation alone does not remove. Ai2’s announcement explains the design.

Read the benchmark conditions alongside the gains

The following are Ai2’s measurements, not independently reproduced results. They answer different systems questions.

Ai2’s reported result

Configuration

What it establishes

4.6B → 47B total parameters; under 5% throughput loss

Eight B300 GPUs; roughly 3.2B active parameters per token; top-4 routing, with 8 → 128 routed experts. Eight-layer BF16 comparison without pipeline parallelism.

A capacity tradeoff in this setup, not a universal price for adding experts.

52,000 versus 19,400 tokens/s/GPU

Preliminary 47B-model comparison on eight B300 GPUs against Ai2’s earlier implementation.

About 2.7× throughput against that baseline, not a ranking of training frameworks.

1.2T total parameters; 858 useful-model TFLOP/s/GPU

512 B300 GPUs; 58.36B active parameters; random routing, MXFP8 and per-layer recomputation.

A selected systems-throughput observation, not a trained-model quality result.

2.38T total-parameter configuration

DeepEP v2; nine steps, without a stabilized throughput plateau.

A short capacity demonstration, not sustained full-run performance.

Do not combine these results into a scaling curve: the larger runs change batch size, topology and precision. Nor do they establish cost to a fixed quality target. The report’s Section 14 and Appendix B.10 make the measurement boundaries explicit.

1. Pin the environment, including optional dependencies

The v3.0.0 package definition requires Python 3.10 or newer and PyTorch ≥2.10.0, <2.14 for the base package. But the fla extra requires PyTorch ≥2.13.0, <2.14, and all includes it. “PyTorch 2.10 minimum” is therefore not a sufficient environment specification for every installation.

Record the release revision, chosen extras, resolved packages, CUDA/driver, attention backend, kernels and communication libraries. Preserve the current environment while qualifying the candidate. The tagged installation README also warns that its dependency images may not suit another cluster’s hardware and drivers.

2. Match a reference recipe to the allocation

The versioned examples assume eight GPUs per node. They need the beaker extra, FlashAttention 4 when using that backend, writable checkpoint/work paths and access to the specified Olmo data mix. DeepEP recipes need a separate checkout. Verify those prerequisites before scheduling GPUs.

The Ultra-128E recipe, OLMoE3-dev-u001.py, specifies 64 nodes, expert-parallel degree 8 and pipeline-parallel degree 8. Its sequence length is 8,192 and global batch is 64 Mi tokens, where Mi means 2²⁰. This is a cluster-scale operating point, not a small-machine quickstart. The example matrix records these settings.

Map expert groups and pipeline stages onto your actual network. Record which communication crosses nodes, together with batch size, expert count, routing mode and precision. NVIDIA’s MoE optimization documentation likewise treats interconnect placement as a performance constraint; it does not supply a matched comparison against Olmo-core.

3. Separate config migration from training continuation

Review these v3.0.0 compatibility changes before relaunching:

  • OLMoDDPModel.apply_ddp() now rejects calls; use apply_dp().

  • A set OLMoDDPTrainModuleConfig.max_grad_norm now overrides the optimizer’s clipping threshold. When unset, the optimizer value remains.

  • The model_ladder API and associated scripts are removed; existing orchestration needs adaptation or an earlier revision.

  • Legacy fused attention loads through FusedAttentionV2 with preserved parameter names and shapes, but the RoPE computation changes. Resumed training is not numerically identical.

For checkpoints with obsolete MoE-v2 class paths, the shipped config migrator rewrites _CLASS_ entries, not model or optimizer tensors. The commands below are upstream-supported examples, not commands executed for this article. From a v3.0.0 checkout with its dependencies installed, substitute a copied checkpoint path and preview first:

python src/scripts/convert_moe_v2_checkpoint_config.py --dry-run /path/to/checkpoint-copy
python src/scripts/convert_moe_v2_checkpoint_config.py --output /path/to/candidate-config.json /path/to/checkpoint-copy

Without --dry-run or --output, the utility rewrites the config in place. A separate output is only a candidate config, not a fully converted checkpoint. Review the changes, then verify restoration of model, optimizer and trainer state in the copied run before changing topology.

Budget storage using total parameters. Section 12 of the report accounts for approximately 12 bytes per parameter for FP32 main weights and two Adam moments. Planning calculation: 12 × 1.2 trillion = 14.4 trillion bytes, or 14.4 decimal TB, before metadata and replication. This is a checkpoint-payload estimate, not a measured Ultra checkpoint or per-GPU memory requirement.

4. Establish a baseline before enabling optimizations

A useful acceptance sequence is to compare the old and candidate stacks with the same checkpoint or initialized weights, data, tokenizer, batches, routing policy, hardware and timing window. Then change precision, recomputation or overlap individually. The checks below are proposed work, not test results:

  • Measure useful tokens per wall-clock second, separating warmup and compilation from the measured window. Include unprofiled end-to-end timing.

  • Track per-rank memory, step-time variation, routing load and dropped-token distributions, plus checkpoint save/restore time.

  • Compare loss and task behavior with learned routing. A random-routing throughput run cannot qualify training quality.

  • Choose acceptance tolerances in advance and retain the old environment and checkpoint for rollback.

The reference README identifies two important limits: the report’s two-batch-overlap experiment cannot be reproduced from this revision, and its supported DDP/FSDP comparison changes routers, optimizers and sparse kernels as well as parallelism. Describe a reproduced result as a comparison of those stacks, not an isolated DDP-versus-FSDP toggle.

5. Validate export separately from training

The tagged Hugging Face export code rejects unsupported combinations, including biased routers or routed experts and non-SwiGLU routed experts. Successful training therefore does not guarantee that every architecture variation can be exported.

Check the actual architecture, tokenizer settings, tensor round-trip and numerical/task behavior in the intended serving runtime. For that downstream step, RohitAI’s vLLM 0.30 migration analysis explains profile-specific acceptance checks; it is not evidence that vLLM supports every Olmo-core export.

For a lab with an active run, the practical next step is a pinned pilot with a copied checkpoint and a representative learned-routing workload. Promote the configuration whose continuation, quality and throughput meet the lab’s requirements—not the one whose parameter count most closely resembles a launch benchmark.

Methodology: AI-assisted reporting and analysis based on Ai2’s announcement, technical report, release metadata and tagged source. No training, package installation, checkpoint conversion or benchmark reproduction was performed. Performance figures remain attributed to Ai2; the storage estimate is explicitly calculated.