0

title: "Mixtral of experts | Mistral AI"

https://mistral.ai/fr/news/mixtral-of-experts/(mistral.ai)
Mistral AI has released Mixtral 8x7B, a high-quality sparse mixture of experts (SMoE) model with open weights. Its architecture uses a router network to select two of eight "expert" parameter groups per token, increasing performance while controlling cost and latency. The model outperforms Llama 2 70B and matches or exceeds GPT-3.5 on most benchmarks with significantly faster inference. Mixtral handles a 32k token context, is multilingual, shows strong performance in code generation, and has an instruction-tuned version available.
0 points•by chrisf•44 minutes ago

Comments (0)

No comments yet. Be the first to comment!

Have an account? Log in to join the discussion.