Mixtral and mixture of experts
Summary
Mixtral activates only part of its experts per token and combines high model capacity with limited computation. The article explains the architecture, usage and open availability. Mixtral uses eight experts and activates two of them per token.
Ideas
- Mixture of experts separates the total number of parameters from the active computing effort.
- Routing becomes an additional model component with its own load and quality questions.
Insights
- Open weights shift control to users and in return require more testing competence of their own.
Facts
- The model weights were made openly available on the Hub.
Critique
- The inactive experts save computing time, but their weights still have to be stored and moved.
Recommendations
- Measure memory requirements, active computing costs and routing behaviour instead of only comparing the number of parameters.
References
Read the original article on Hugging Face
Links to the original source and the Web Archive open in a new tab.