Digital Event Horizon
Revolutionizing Large Language Model Training: Olmo-core 3 Unveiled
Hugging Face introduces Olmo-core 3, a cutting-edge training infrastructure for large MoE models. Olmo-core 3 is designed to handle trillion-parameter capacities, a feat previously deemed impossible. The framework is built on key techniques for distributing large MoEs across GPU clusters with optimizations for routing and computation. Olmo-core 3 reduces the cost of routing data to the right experts and running their computations. The framework supports MXFP8, a lower-precision number format, to reduce computation and data movement. Olmo-core 3 has been benchmarked on various configurations, achieving high throughput and scalability. The framework is fully open, allowing researchers and developers to use, adapt, and experiment with it.
Hugging Face has made a significant breakthrough in the field of large language model training with the introduction of Olmo-core 3, a cutting-edge training infrastructure designed to scale the training of large MoE (Mixture of Experts) models. This innovative framework is built to handle the computational demands of training models with trillion-parameter capacities, a feat previously deemed impossible.
The development of Olmo-core 3 is the culmination of a decade-long effort by Hugging Face to build a training stack that accurately models how MoEs work. The company's researchers have been working tirelessly to optimize the training process, with a primary focus on scaling and optimizing MoE training. This new infrastructure is a significant upgrade to the framework behind Olmo, which is one of the core systems behind the next generation of Olmo.
Training large AI models has become increasingly challenging due to the enormous computational requirements involved. MoE models offer a more efficient approach, as they can contain many more learned components without requiring every input to use all of them. However, the full model still needs to be stored across GPU memory and updated during training, and directing inputs to the right experts across a cluster creates its own communication and coordination costs. As MoEs grow, those costs can erode much of the computational advantage of using only part of the model for each input.
Olmo-core 3 is designed to close this gap. In a benchmark, the new framework increased the expert pool from 8 to 128 while still selecting only four experts per token, keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%. Furthermore, Olmo-core 3 has been benchmarked at over one trillion total parameters.
The development of Olmo-core 3 is built on the back of several key techniques for distributing large MoEs across GPU clusters with optimizations that make routing and computation more efficient. Three techniques determine how the model and its training state are split across hardware: expert parallelism, pipeline parallelism, and a distributed optimizer. These techniques allow an MoE to scale without requiring every GPU to keep the entire model and its training state in memory.
In addition, Olmo-core 3 reduces the cost of routing data to the right experts and running their computations. Techniques such as rowwise expert parallelism, GPU-resident routing, and grouped GEMM combine many small expert computations, allowing GPUs to execute them more efficiently. The framework also supports MXFP8, a lower-precision number format that represents some values with fewer bits, which can reduce computation and the amount of data moved between GPUs.
Olmo-core 3 has been benchmarked on a range of configurations on NVIDIA B300 GPUs, including a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs. Its highest observed throughput was 858 TFLOP/s/GPU—a measure of useful model computation per second on each GPU. These tests used random routing to measure system performance, rather than the quality of a trained model.
The development of Olmo-core 3 is part of Hugging Face's broader commitment to open up the tools and training infrastructure behind each new model. The framework is designed to be fully open, allowing researchers and developers to use Olmo-core 3 to train their own MoEs, adapt it to different hardware, and experiment with routing, parallelism, and other parts of the system.
The next-generation Olmo will use an MoE architecture, and the company is aiming for it to be its most capable Olmo yet, trained on its largest dataset and with its longest context window. Olmo-core 3 provides the necessary infrastructure to scale beyond previous MoE work while giving researchers and developers more flexibility to adapt training as models and hardware evolve.
In conclusion, Olmo-core 3 represents a major breakthrough in large language model training, providing a scalable and efficient training infrastructure for MoE models. The framework is designed to handle the computational demands of training models with trillion-parameter capacities, and its optimizations and techniques make it an ideal solution for researchers and developers looking to train large MoEs.
Related Information:
https://www.digitaleventhorizon.com/articles/Revolutionizing-Large-Language-Model-Training-Olmo-core-3-Unveiled-deh.shtml
https://huggingface.co/blog/allenai/olmocore3
Published: Thu Oct 1 10:23:30 2026 by llama3.2 3B Q4_K_M