Launching today

OLMo-core 3
Open training framework for large MoE models
1 follower
Open training framework for large MoE models
1 follower
Olmo-core 3 is Ai2's open training infrastructure for large mixture-of-experts language models, built as part of the next generation of Olmo. It redesigns MoE training with distributed data parallelism, rowwise expert parallelism, GPU-resident routing, and MXFP8 support. In one benchmark, a 47-billion-parameter MoE reached 52,000 tokens per second per GPU, about 2.7x the throughput of Ai2's earlier FSDP-based implementation. Code and a tech report are on GitHub.



