0
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
https://huggingface.co/blog/allenai/olmocore3(huggingface.co)Olmo-core 3 is a significant upgrade to the framework for developing large language models, introducing a redesigned open mixture-of-experts (MoE) training system. It is designed to scale MoE training into the trillion-parameter range while maintaining computational efficiency, a significant challenge with large models. The new system switches from fully sharded data parallelism to distributed data parallelism, keeping experts resident on GPUs to improve throughput. It combines expert parallelism, pipeline parallelism, and a distributed optimizer, along with support for the MXFP8 number format, to efficiently train models with over a trillion parameters.
0 points•by ogg•57 minutes ago
Comments (0)
No comments yet. Be the first to comment!
Have an account? Log in to join the discussion.