Launching today
OLMo-core 3

OLMo-core 3

Open training framework for large MoE models

1 follower

Olmo-core 3 is Ai2's open training infrastructure for large mixture-of-experts language models, built as part of the next generation of Olmo. It redesigns MoE training with distributed data parallelism, rowwise expert parallelism, GPU-resident routing, and MXFP8 support. In one benchmark, a 47-billion-parameter MoE reached 52,000 tokens per second per GPU, about 2.7x the throughput of Ai2's earlier FSDP-based implementation. Code and a tech report are on GitHub.

No makers yet

It looks like there are no makers for this product.